James Cheng
@alexander65
Senior software engineer building scalable AI infrastructure, distributed systems, and low-latency machine learning platforms.
What I'm looking for
I’m a Senior Software Engineer with 10+ years of experience building scalable distributed systems, AI infrastructure, and high-performance backend platforms across cloud-native environments. I focus on production-grade machine learning infrastructure that delivers reliability, speed, and cost efficiency.
At Perplexity AI, I architected and scaled a production RAG platform handling 25M+ daily retrieval queries with sub-second latency. I built high-throughput ingestion pipelines, integrated hybrid vector retrieval (FAISS and Pinecone) to improve grounding accuracy, and developed modular LLM orchestration with multi-provider support. I optimized GPU-backed inference with batching and autoscaling, implemented caching to cut token costs, and established observability and resilience practices that supported 99.99% availability.
Previously, at Ramp and Confluent, I led low-latency ML inference APIs and Kafka-based event streaming systems, strengthening exactly-once guarantees, uptime, and operational readiness. Earlier at Elastic, I worked on distributed search and indexing performance for large-scale datasets while improving reliability through validation, logging, and automated testing.
Experience
Work history, roles, and key accomplishments
Architected and scaled a production RAG platform processing 25M+ daily retrieval queries with sub-second latency across multi-provider LLM environments. Built high-throughput ingestion and observability, and implemented GPU inference optimizations and resilience mechanisms to maintain 99.99% availability.
Led development of scalable ML inference APIs for real-time expense classification and fraud detection across 120M+ annual financial transactions. Built low-latency feature retrieval and Kafka-based streaming pipelines, and productionized models into Kubernetes microservices.
Developed backend infrastructure services for provisioning and managing thousands of enterprise Kafka clusters handling petabyte-scale streaming workloads. Built REST and gRPC APIs for topic lifecycle management and improved reliability through exactly-once processing and observability tooling.
Supported development of large-scale Kafka distributed systems and backend infrastructure for enterprise streaming environments. Contributed to improvements in cluster stability, exactly-once processing reliability, and event-driven systems practices.
Developed backend services for Elasticsearch indexing and distributed search APIs for enterprise customers managing multi-terabyte datasets. Optimized shard allocation and indexing configurations, built scalable ingestion pipelines, and improved reliability through testing and automated error handling.
Education
Degrees, certifications, and relevant coursework
Liberty University
Bachelor of Science, Computer Science
2011 - 2015
Bachelor of Science in Computer Science at Liberty University from 2011 to 2015.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring James?
You can contact James and 90k+ other talented remote workers on Himalayas.
Message JamesGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
