At LTM, I build production ETL pipelines that process 20–30 million daily records, migrating Hadoop HDFS data to Amazon S3 while maintaining approximately 98.5% pipeline uptime.
I developed 10+ end-to-end PySpark and AWS ingestion pipelines for banking data in CSV, JSON, and Parquet formats, using S3, EMR, and Glue to onboard sources into an enterprise data lake. I also built curated data marts for customer e-Statement generation and fraud analysis, reducing report preparation time and supporting earlier identification of suspicious transactions.
I improve data quality through Python-based validation frameworks, raising accuracy from 85% to 92% across five key data sources. My work includes SQL optimization, transformation, cleansing, validation, and reliable issue tracking.
My foundation includes hands-on AWS, Linux, Hadoop/HDFS, Hive, Airflow, PySpark, Python, and advanced SQL, developed through data engineering training and production delivery.

