At Infosys on the Global Foundries account, I build AWS-based data platforms that turn multi-source raw data into governed, analytics-ready Redshift data marts. I designed a JSON-driven PySpark ETL framework across 70+ tables, reducing Power BI reporting latency by 70% for five-plus business departments.
I've improved data accuracy to 99.5% across 10+ TB of enterprise data through automated validation, monitoring, source-to-target mapping, and metadata documentation. My work includes dimensional modeling with star schemas and SCD Type 2, full and incremental load logic, and Apache Iceberg on Amazon S3 to improve multi-terabyte incremental update performance by 65%.
I also conduct peer reviews, standardize reusable pipeline patterns, and mentor junior engineers on data modeling and validation practices.
