At Ember AI, I architected a production litigation platform that orchestrates roughly 40 specialized LLM agents through seven staged synthesis pipelines for paying users.
I reduced token spend by roughly 60% through Anthropic prompt caching, dossier memoization, and async-bounded retrieval. I also built agentic RAG systems for legal document analysis, combining Gemini vision OCR, Claude reasoning, Pinecone retrieval, and service-oriented FastAPI backends.
I shipped Artie, a live AI-powered media platform, end to end with FastAPI, Celery/Redis, React, OpenAI, Gemini, and AWS Amplify. My public Ask FastAPI Docs RAG demo captures the production patterns I use in client work, including hybrid retrieval, structured SSE streaming, and evaluation workflows.
Previously, I built async web applications at Devsinc and developed Flask and Django applications during my internship at Conovo Technologies. I use Cursor and Claude Code daily, working from markdown specs and keeping engineers in control of AI-assisted implementation.
