Skip to main content
SM
Looking for a job

Sai Mahanth

@saimahanth

I build cloud data pipelines, lakehouse models, and reliable analytics from high-volume data.

Germany
Message

What I'm looking for

I'm looking to build reliable cloud data platforms, real-time pipelines, and analytics-ready models while collaborating closely with data scientists and cross-functional teams.

I've built production data pipelines at Hindustan Aeronautics Limited and AI ML Labs, turning high-volume simulation and user-activity data into reliable analytics and ML-ready datasets.

At HAL, I engineered an aerospace pipeline for 2M+ simulation records and created a PySpark validation framework that removed 35% of duplicates, improving model training accuracy by 18%.

At AI ML Labs, I architected a real-time Python and Apache Kafka pipeline processing 50K+ events daily, while building ETL/ELT workflows with Azure Data Factory and Airflow that reduced downstream errors by 40%.

I'm focused on Azure Databricks, Delta Lake, SQL, Python, dimensional modeling, and data quality. My recent projects span Snowflake dbt ELT, healthcare ETL, real-time Uber ride analytics, RAG systems, and petroleum trade forecasting.

Experience

Work history, roles, and key accomplishments

PP

Airbnb dbt project

Personal Project

Sep 2026 - Sep 2026 (0 months)

Built an end-to-end ELT pipeline on Snowflake using dbt, implementing a medallion architecture (bronze/silver/gold) to transform raw Airbnb data into analytics-ready models. Designed fact tables and dimensional models in the gold layer, with SCD Type 2 history tracking via dbt snapshots.

PP

Agentic Resume-Screening RAG Pipeline

Personal Project

Aug 2026 - Aug 2026 (0 months)

Built and evaluated a multi-stage RAG pipeline over 25 resumes (LangChain, OpenAI API); ran systematic error analysis to diagnose retrieval failure modes, then benchmarked fixes that improved retrieval precision by an estimated 30%. Implemented structured, grounded output with Pydantic schemas and validation checks that reduced hallucinated answers by an estimated 40%.

PP

Healthcare Revenue Cycle Management

Personal Project

Jun 2026 - Jun 2026 (0 months)

Built parameterized ETL pipelines in Azure Data Factory ingesting EMR data from Azure SQL to ADLS Gen2 Bronze/Silver/Gold layers (5+ hospital sources); implemented incremental load strategies with watermark triggers reducing processing time by 40%. Developed PySpark notebooks for data cleansing and standardization of medical codes (ICD-10, CPT, HCPCS).

PP

Real-Time Streaming Pipeline

Personal Project

May 2026 - May 2026 (0 months)

Built end-to-end streaming pipeline using Azure Event Hubs ingesting ride events into Databricks Lakeflow with bulk loads into ADLS Gen2; architected Medallion Lakehouse supporting high-volume real-time ingestion, processed 15,000+ ride events per day with sub-2-minute end-to-end latency. Designed Gold-layer star schema using Databricks CDC flows with SCD Type 1/2 dimensions.

US

Master Thesis: Global Refined Petroleum Trade Forecasting Analysis

University of Europe for Applied Sciences

Nov 2025 - Feb 2026 (3 months)

Designed scalable Azure Data Factory ETL pipelines to process 15+ years of UN Comtrade petroleum trade data across 100+ countries; applied forecasting and network analysis, achieving an estimated 85% forecast accuracy (MAPE under 15%) in modeling regional trade flows through 2028.

Hindustan Aeronautics Limited logoHL

Software Engineer

Hindustan Aeronautics Limited

Apr 2023 - Dec 2023 (8 months)

Engineered aerospace data pipeline ingesting 2M+ simulation records; removed 35% duplicate entries through custom PySpark validation framework, improving model training accuracy by 18%. Designed dimensional data models and schemas enabling 3 cross-functional teams to run performance benchmarking analysis independently.

AL

Data Engineer

AI ML Labs Pvt. Ltd.

Feb 2022 - Feb 2023 (1 year)

Architected real-time user activity pipeline (Python, Apache Kafka) ingesting 50K+ events/day; dual-storage strategy enabling analytics on 6M+ event records with sub-100ms query latency. Built 5+ production ETL/ELT pipelines (Azure Data Factory, Airflow) implementing Medallion architecture and data validation checks, reducing downstream errors by 40%.

Education

Degrees, certifications, and relevant coursework

University of Europe for Applied Sciences logoUS

University of Europe for Applied Sciences

Master of Science, Data Science

2024 - 2026

Pursued a Master of Science in Data Science, focusing on data engineering and analytics.

Visvesvaraya Technological University logoVU

Visvesvaraya Technological University

Bachelor of Engineering, Mechanical Engineering

2017 - 2021

Earned a Bachelor of Engineering in Mechanical Engineering.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan