I've built large-scale distributed data processing applications and end-to-end ETL pipelines using Apache Spark, PySpark, Spark SQL, AWS EMR, S3, and Apache Airflow.
At Zunax Energy LLP, I design batch analytics workflows for structured and semi-structured data, applying partitioning, caching, persistence, shuffle tuning, and memory optimization to improve Spark job performance and resource utilization. I also implement data cleansing, validation, preprocessing, transformations, aggregations, and joins using Spark RDDs and DataFrames.
Previously at Greenmetro Group, I developed scalable Spark workflows, integrated Spark with Hadoop and Hive, monitored production jobs, and resolved data-processing bottlenecks. I bring over five years of hands-on experience with distributed processing, data lakes, workflow orchestration, and reliable ETL delivery.
