At Sigmoid Analytics, I designed Databricks and PySpark pipelines for a GenAI-powered document retrieval platform, processing more than 500K structured and unstructured documents. The platform reduced manpower requirements by 60% and shortened request turnaround time by removing manual technical handoffs.
I implemented a Bronze–Silver–Gold Medallion architecture on Delta Lake and built incremental ingestion from source APIs into ADLS. I also developed document processing workflows from extraction and OCR through chunking and embedding generation, publishing results to Azure AI Search for vector and semantic retrieval.
Earlier at Sigmoid Analytics, I maintained ETL/ELT pipelines and optimized Spark SQL transformations, reducing incremental pipeline runtime by 70%. I also built data quality frameworks that reduced incorrect or duplicate ingestion by over 90%, and completed a Databricks training capstone using a Medallion architecture and star schema.
