Skip to main content
HR
Open to opportunities

Harshita Ragha

@harshitaragha

I build scalable AI data pipelines and production machine learning platforms.

United States
Message

What I'm looking for

I'm looking to build scalable AI data infrastructure and production ML platforms, applying distributed processing, dataset quality and lineage, cloud architecture, and RAG systems to reliable enterprise AI applications.

At Comerica Bank, I build distributed AI data pipelines and enterprise RAG applications across GCP and Azure. I improved vector-search query latency by 28% and maintained 99.9% production uptime through Docker, logging, and CI/CD workflows.

Previously at Banner Health, I developed secure ingestion and distributed processing services handling more than 7M customer and operational records daily. I reduced batch-processing time by 30% through SQL, indexing, and schema optimization.

My background spans Python, PySpark, Spark, Kafka, dataset lineage, data quality, and scalable backend systems, with experience delivering high-volume platforms from Loopbell Tech and iConcept Software Services to healthcare and banking environments.

Experience

Work history, roles, and key accomplishments

Comerica Bank logoCB
Current

Gen AI/ML Engineer

Comerica Bank

Nov 2024 - Present (1 year 9 months)

Built distributed data pipelines with Python, PySpark, Apache Spark, and Kafka to ingest, transform, validate, and prepare high-volume enterprise datasets for AI training and evaluation. Tuned dataset storage and retrieval layouts to reduce I/O bottlenecks and developed enterprise RAG applications with vector search, reducing query latency by 28%.

Banner Health logoBH

Gen AI/ML Engineer

Aug 2023 - Oct 2024 (1 year 2 months)

Developed backend data ingestion and distributed processing services using Python, Java, and PySpark ETL pipelines to securely process over 7M+ records daily. Refactored SQL queries and database schemas, leading to a 30% drop in batch-processing time and maintaining 99.4% system reliability.

IL

Data Engineer

iConcept Software Services Pvt. Ltd.

Jan 2020 - Jul 2021 (1 year 6 months)

Engineered scalable backend data ingestion and processing services using Python, Java, and PySpark ETL pipelines, securely handling over 7M+ records daily. Optimized SQL queries and database schemas, achieving a 30% reduction in batch-processing time and maintaining 99.4% system availability.

LL

Data Engineer

Loopbell Tech Private Limited

Jun 2018 - Dec 2019 (1 year 6 months)

Built backend data ingestion services and RESTful APIs using Python, SQL, and Pandas to handle 5M+ daily event logs. Optimized relational database structures in MySQL and PostgreSQL, lowering critical query latency by 15% and streamlining distributed data transformation jobs with PySpark.

Education

Degrees, certifications, and relevant coursework

University of North Texas logoUT

University of North Texas

Master of Science, Data Science

2021 - 2023

Pursued a Master's in Data Science, focusing on advanced data analytics and machine learning techniques.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan