At Vidvanconnect Private Limited, I build reusable PySpark and SQL workflows for migrating Curated News from Azure Data Factory to Databricks. I added watermarks, deduplication, reconciliation, and invalid-record quarantine to support reliable incremental processing.
I also refactored document ingestion into modular FastAPI services and built secured APIs for uploads, search, and audit history. These workflows integrate Azure Blob Storage, MongoDB/Cosmos DB, and Azure Document Intelligence for document and web-content ingestion.
At Radisys Private Limited, I processed and transformed datasets totaling 600+ TB for JioCinema IPL analytics using Spark DataFrames, Scala, Hive, and Spark SQL. I optimized distributed workloads and built batch and streaming retail pipelines with Kafka, Spark Structured Streaming, and Delta Lake.
Earlier, at Cognizant Technology Solutions, I developed Spark and Scala workflows on Amazon EMR, stored transformed data in Amazon S3, and built incremental ingestion from RDBMS sources into HDFS and Hive. I also tuned analytical transformations and orchestrated workflows with Oozie.

