At Vellum AI, I architect and deliver enterprise-grade generative AI systems using LLMs, RAG, and agent-based architectures. I reduced response latency by 35% through efficient retrieval strategies and optimized data pipelines.
I established prompt engineering frameworks and evaluation methods that increased LLM response accuracy by 30%. I also engineered cloud deployment pipelines with AWS, Docker, and CI/CD automation, shortening release cycles by 40%.
Previously, at ElevenLabs, I developed real-time inference pipelines that reduced processing latency by 40% and enhanced voice generation quality by 25%. At Labelbox, I developed machine learning workflows for data preparation, annotation, and model training across computer vision and NLP applications.

