I've built Azure data pipelines at Quantix Innovations Pvt Ltd, consolidating banking transaction and account data from six regional SQL Server databases into a centralized warehouse processing 800K daily records.
I design parameterized Azure Data Factory pipelines and Databricks PySpark transformations across Bronze, Silver, and Gold Delta Lake layers. By replacing full-load ingestion with watermark-based incremental pipelines, I reduced daily data movement from 2.3 GB to 800 MB.
I implement SCD Type 2 logic with Delta Lake MERGE INTO, preserving historical dimension data while keeping current records fast for reporting. I also validate nulls, duplicates, schema consistency, and source-to-gold reconciliation before data reaches reporting layers.
I've supported retail production pipelines, resolved schema drift and job failures through root-cause analysis, and tuned PySpark partitions and cluster configurations to keep batch workloads within SLA. I hold the Microsoft DP-203 Azure Data Engineer Associate certification.
