I've built production LLM and distributed systems at Interface AI, including a multi-provider agent runtime with tool calling, context management, prompt caching, retries, and per-run cost accounting.
I built an evaluation harness across 130+ tools and 300+ scenarios, and model routing that reduced inference spend by about 27% while maintaining equal evaluation pass rates. I also architected a multi-tenant sourcing engine for 10M+ profiles and led a zero-downtime migration from GCP to AWS ECS Fargate.
Previously, I delivered real-time speech, RAG, monitoring, IAM, MFA, and SSO systems, from 4K+ concurrent ASR streams to Shopify and WordPress 2FA integrations.
