Skip to main content
Scott ChenSC
Open to opportunities

Scott Chen

@scottchen5

I build LLM retrieval and real-time data infrastructure at production scale.

United States
Message

At Scale AI, I build the ingestion, validation, retrieval, and evaluation layers behind production LLM applications. I process roughly 40 TB per month of multimodal customer data and reduced re-labeling from bad source data by approximately 35%.

I’ve built embedding pipelines, pgvector and Milvus indexes, and hybrid RAG retrieval using BM25, dense search, reciprocal rank fusion, and cross-encoder reranking. This improved recall@10 from 0.71 to 0.89 while reducing ungrounded responses, and my telemetry-driven prompt caching, batching, and model routing helped reduce inference spend by approximately 30%.

Previously at DoorDash, I helped modernize real-time event ingestion with Kafka and Flink on Kubernetes, reducing analytics latency from roughly 24 hours to minutes across 200+ event tables. I also built systems handling billions of events per day at 99.99% delivery and mentored three junior engineers.

My earlier work at Optum and EPAM Systems grounded me in governed healthcare data, data quality, identity resolution, analytics, and business intelligence. I build production systems end to end with strong attention to reproducibility, observability, schema evolution, and secure data handling.

Experience

Work history, roles, and key accomplishments

Education

Degrees, certifications, and relevant coursework

Texas State University logoTU

Texas State University

Bachelor of Science, Computer Science

2012 - 2016

Bachelor of Science in Computer Science from Texas State University, completed from 2012 to 2016.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan