Skip to main content
V S S Mrunalini PenmatchaVP
Open to opportunities

V S S Mrunalini Penmatcha

@vssmrunalinipenmatch

Data Engineer focused on scalable AWS data platforms, performant ETL/ELT, and reliable analytics.

United States
Message

What I'm looking for

I’m looking for a data engineering role where I can build scalable AWS ETL/ELT pipelines, improve data quality and performance, and automate workflows—so teams get reliable, self-service analytics and faster time to insight.

I’m a data professional with 6.8 years of experience building scalable data platforms, distributed ETL/ELT pipelines, and cloud-native data solutions on AWS. I specialize in SQL, Python, Spark, PySpark, and data warehousing and lakes, with a strong focus on performance optimization and automation.

At Capital One, I designed and optimized ETL pipelines using PySpark and AWS Glue for large-scale financial datasets, delivering a 35–45% reduction in processing time and improved downstream data availability. I also led enterprise data pipelines for regulatory reporting, improving data accuracy by nearly 30% while reducing reconciliation issues, and raised test coverage from 50% to 97% using PyTest automation to reduce production defects.

In my capstone work with Asurion, I built an automated label parsing workflow using OpenCV and Pytesseract OCR, orchestrated with AWS Step Functions and AWS Lambda. Using AWS EMR for distributed processing, I improved image quality by 30% and achieved 85% accuracy in serial number extraction via a machine learning approach.

Earlier, as a Software Engineer at Tech Mahindra (GE Healthcare), I supported asset performance management by configuring Predix APM and developing Spark MLlib pipelines for predictive maintenance, improving failure prediction accuracy by 25% and reducing downtime. I’ve consistently applied software engineering best practices—CI/CD, unit testing, code reviews, and Agile delivery—while building analytics-ready datasets and self-service BI layers.

Experience

Work history, roles, and key accomplishments

Capital One logoCO

Data Engineer

Nov 2023 - Dec 2025 (2 years 1 month)

Designed and optimized ETL pipelines using PySpark and AWS Glue for large-scale financial datasets, reducing processing time by 35–45% and improving downstream data availability. Built enterprise data pipelines for regulatory reporting, improved data accuracy by nearly 30%, and strengthened production reliability with monitoring (CloudWatch/Splunk) and CI/CD best practices.

Tech Mahindra Ltd. logoTL

Software Engineer - APM

Mar 2018 - Aug 2021 (3 years 5 months)

Supported Asset Performance Management (APM) by configuring a Predix customer tenant and managing data ingestion for analytics. Developed predictive maintenance machine learning pipelines with Spark MLlib and optimized SQL/SparkSQL queries, improving failure prediction accuracy by 25% and reducing query execution time by 35–40%.

Talent Sprint logoTS

Data Engineer

Talent Sprint

Jan 2016 - Jun 2016 (5 months)

Used Hadoop (HDFS and MapReduce) to store, retrieve, and preprocess large population datasets for predictive analysis and anomaly detection. Implemented ML approaches (linear regression, logistic regression, clustering), reducing data processing time by 40% and improving population trend forecast precision by 25%.

Education

Degrees, certifications, and relevant coursework

George Mason University logoGU

George Mason University

Master's in Data Analytics Engineering, Data Analytics Engineering

Activities and societies: Certifications: Google Data Analytics Certificate; HackerRank SQL Basic/Intermediate/Advanced; HackerRank Problem Solving (Basic).

Earned a Master's in Data Analytics Engineering at George Mason University in Fairfax, VA.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan