Skip to main content
YM
Open to opportunities

Yashwanth Goud Matta

@yashwanthgoudmatta

I build high-scale AI infrastructure for LLM inference, RLHF, and distributed GPU systems.

United States
Message

What I'm looking for

I'm looking to build production AI platforms where I can scale LLM inference, RLHF, GPU infrastructure, and event-driven systems while improving reliability, safety, and developer or customer outcomes.

At Scale AI, I architect enterprise RLHF and human-feedback systems processing 15B+ monthly inference requests, reducing inference costs by 29% while improving LLM alignment quality by 34%. I also build governance, multimodal processing, personalization, and workflow platforms using Java, Spring Boot, Kafka, Kubernetes, and distributed GPU infrastructure.

Previously at NVIDIA and Meta, I built GPU-orchestrated training and inference platforms, real-time control planes, and high-throughput ads ranking services. My work spans low-latency microservices, event-driven systems, Kubernetes, RAG, model serving, and production AI observability.

Experience

Work history, roles, and key accomplishments

Scale AI logoSA
Current

AI/ML Engineer

Mar 2025 - Present (1 year 5 months)

Architected enterprise RLHF platform processing 15B+ monthly inference requests, reducing inference costs by 29% and improving LLM alignment quality by 34%. Engineered low-latency AI orchestration services achieving p95 latency under 350ms and built human-in-the-loop annotation systems supporting 420K+ reviewers.

Education

Degrees, certifications, and relevant coursework

UD

University of Colorado Denver

Master of Science, Business Analytics

Pursued a Master's in Business Analytics at the University of Colorado Denver.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan