At Infosys, I've designed and developed Spark applications for complex batch transformations and aggregations, supporting the Broadcom project. I've built data lake architectures on Amazon S3 using partitioning and Parquet to improve query performance and reduce storage costs.
I've optimized PySpark and Spark RDD workloads on AWS EMR, orchestrated workflows with AWS Step Functions and Airflow, and delivered large-scale Sqoop migrations from on-premise databases to Hadoop clusters. I also troubleshoot performance, memory, scalability, and data-processing issues across Spark, Hive, and Sqoop environments.
