At A&N SOFTWARE SOLUTIONS, I work on the Capital One project as a Data Engineer, developing Spark applications for distributed data processing. I build batch and streaming pipelines and use Spark transformations to cleanse, preprocess, and enrich data.
I orchestrate PySpark workflows on AWS EMR with Apache Airflow, configuring EMR steps and cluster resources for workload requirements. I’ve also automated EMR cluster lifecycle management using Airflow operators and sensors.
I optimize Spark jobs through partitioning, caching, and resource tuning, and troubleshoot performance bottlenecks. My work includes integrating Spark with Hive and HBase, and using formats such as Avro, Parquet, and JSON.

