I'm building end-to-end data pipelines at TCS using Azure Data Factory and Databricks, managing daily ingestion of structured and semi-structured data.
I ingest data from on-premise Oracle DB and SQL Server into ADLS Gen2, and write complex transformation logic in PySpark. I've reduced Spark job execution time through Z-Ordering, partitioning, and broadcast joins.
I've contributed to Medallion Architecture implementations across Bronze, Silver, and Gold layers, improving data reliability and downstream analytics performance. I also collaborate with stakeholders to understand requirements and build solutions in the pharma domain.
Previously, I developed and maintained Informatica PowerCenter ETL pipelines, optimized SQL queries and indexing strategies, and helped design a centralized data warehouse serving as a single source of truth for BI and reporting applications.
