V S S Mrunalini Penmatcha
@vssmrunalinipenmatch
Data Engineer focused on scalable AWS data platforms, performant ETL/ELT, and reliable analytics.
What I'm looking for
I’m a data professional with 6.8 years of experience building scalable data platforms, distributed ETL/ELT pipelines, and cloud-native data solutions on AWS. I specialize in SQL, Python, Spark, PySpark, and data warehousing and lakes, with a strong focus on performance optimization and automation.
At Capital One, I designed and optimized ETL pipelines using PySpark and AWS Glue for large-scale financial datasets, delivering a 35–45% reduction in processing time and improved downstream data availability. I also led enterprise data pipelines for regulatory reporting, improving data accuracy by nearly 30% while reducing reconciliation issues, and raised test coverage from 50% to 97% using PyTest automation to reduce production defects.
In my capstone work with Asurion, I built an automated label parsing workflow using OpenCV and Pytesseract OCR, orchestrated with AWS Step Functions and AWS Lambda. Using AWS EMR for distributed processing, I improved image quality by 30% and achieved 85% accuracy in serial number extraction via a machine learning approach.
Earlier, as a Software Engineer at Tech Mahindra (GE Healthcare), I supported asset performance management by configuring Predix APM and developing Spark MLlib pipelines for predictive maintenance, improving failure prediction accuracy by 25% and reducing downtime. I’ve consistently applied software engineering best practices—CI/CD, unit testing, code reviews, and Agile delivery—while building analytics-ready datasets and self-service BI layers.
Experience
Work history, roles, and key accomplishments
Designed and optimized ETL pipelines using PySpark and AWS Glue for large-scale financial datasets, reducing processing time by 35–45% and improving downstream data availability. Built enterprise data pipelines for regulatory reporting, improved data accuracy by nearly 30%, and strengthened production reliability with monitoring (CloudWatch/Splunk) and CI/CD best practices.
Built a serverless image preprocessing and OCR workflow using OpenCV and AWS orchestration to extract machine-readable text and improve serial number extraction. Orchestrated preprocessing with AWS Step Functions/Lambda and used OCR via PyTesseract, achieving 85% accuracy for serial number extraction.
Supported Asset Performance Management (APM) by configuring a Predix customer tenant and managing data ingestion for analytics. Developed predictive maintenance machine learning pipelines with Spark MLlib and optimized SQL/SparkSQL queries, improving failure prediction accuracy by 25% and reducing query execution time by 35–40%.
Data Engineer
Talent Sprint
Jan 2016 - Jun 2016 (5 months)
Used Hadoop (HDFS and MapReduce) to store, retrieve, and preprocess large population datasets for predictive analysis and anomaly detection. Implemented ML approaches (linear regression, logistic regression, clustering), reducing data processing time by 40% and improving population trend forecast precision by 25%.
Education
Degrees, certifications, and relevant coursework
George Mason University
Master's in Data Analytics Engineering, Data Analytics Engineering
Activities and societies: Certifications: Google Data Analytics Certificate; HackerRank SQL Basic/Intermediate/Advanced; HackerRank Problem Solving (Basic).
Earned a Master's in Data Analytics Engineering at George Mason University in Fairfax, VA.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring V S S Mrunalini?
You can contact V S S Mrunalini and 90k+ other talented remote workers on Himalayas.
Message V S S MrunaliniGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
