Maithil Deore
@maithildeore
Data Engineer specializing in batch/streaming pipelines with lakehouse and quality-first ETL.
What I'm looking for
I’m a Data Engineer with 3+ years of production experience designing and optimizing batch and streaming pipelines across 100 TB+ Hadoop/HDFS and Spark platforms. I’ve cut SLA-breaching PySpark job runtimes by 25% and helped drive $75K+/yr in infrastructure savings.
I build modern lakehouse pipelines with Databricks and Delta Lake, using medallion (Bronze/Silver/Gold) architecture and production-ready streaming patterns. I’ve delivered Kafka-driven ingestion and transformation pipelines supporting 500K+ subscribers at 5M+ records/month.
My work is defined by data-quality discipline: I integrate schema validation, row-count reconciliation, and Great Expectations checks at every pipeline layer. I also build data-access APIs with FastAPI/REST to bridge data engineering and data product delivery.
As AWS-primary, I leverage S3, EMR, Glue, Lambda, Athena, Kinesis, and orchestration patterns to ship reliable systems, while bringing working Azure exposure. I’m currently AWS Certified Data Engineer (Associate) focused and in Databricks Data Engineer certification progress.
Experience
Work history, roles, and key accomplishments
Data Engineer / Backend Developer
Businessnext
Oct 2024 - Present (1 year 9 months)
Translated business requirements into ETL components and added schema-validation and row-count reconciliation to prevent upstream data corruption reaching BI dashboards. Developed FastAPI-based REST data-access APIs and introduced Great Expectations checks to reduce data-quality incidents.
Data Engineer
Kyndryl
Aug 2023 - Sep 2024 (1 year 1 month)
Diagnosed and rewrote SLA-breaching PySpark jobs to improve performance (partition-pruned reads, broadcast joins, predicate pushdown) and reduce runtime and CPU-hours. Built end-to-end PySpark pipelines and restructured Hive tables to ORC with partitioning/bucketing to significantly improve analytics query latency.
Data Science Intern
TCR Innovation
Oct 2021 - Mar 2022 (5 months)
Built machine learning classification models for HR attrition and heart-disease prediction using feature engineering and cross-validation. Worked with Python-based data science libraries to train and evaluate models.
Education
Degrees, certifications, and relevant coursework
Savitribai Phule Pune University
Bachelor of Engineering (Information Technology), Information Technology
2019 - 2023
Grade: 9.0 CGPA (Data Science Honours); 8.0 CGPA (B.E. IT)
Earned a B.E. in Information Technology with Data Science Honours, achieving a 9.0 CGPA and publishing an academic work on disease prediction using lab reports.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Maithil?
You can contact Maithil and 90k+ other talented remote workers on Himalayas.
Message MaithilGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
