At REGex Software Services, I build Python and PySpark pipelines in Databricks to process data from AWS S3 for downstream analytics. I design incremental S3-to-Snowflake workflows and implement Bronze-Silver-Gold layers with validation checks, delivering analytics-ready datasets at 99%+ data quality.
On my Flight Delay Analytics project, I processed 2M+ records with PySpark on Databricks, using partition pruning and caching to cut processing time by roughly 45%. I also built a Google Drive-to-AWS pipeline that removed more than two hours of manual uploads each day, and developed SQL queries at Celebal Technologies that reduced report-generation time by nearly 40%.

