At MERIL (Nuvo AI), I built an asynchronous RAG orchestration platform with FastAPI and asyncio, using RabbitMQ workers and Qdrant to handle retrieval workloads outside the request path. I also developed a biomedical document-processing pipeline that achieved a 10x processing speedup through dynamic chunking and parallel inference.
I deployed and served LLMs on on-premise GPU infrastructure with vLLM, and containerized services with Docker and Kubernetes. I improved RAG retrieval precision from 71% to 87% and reduced query latency from 450ms to 180ms through context-aware chunking, MMR reranking, and Qdrant indexing optimizations.
I designed a stateful healthcare copilot with LangGraph and custom MCP servers for healthcare integrations, including EHR APIs and knowledge graph queries. Its workflows cover intent planning, entity resolution, tool execution, and conversation state, with Redis-backed memory and reflexion workflows for context retention and hallucination control.
I also built an LLM-as-a-Judge evaluation framework that reached 94% accuracy for protocol-deviation detection against GPT-4 and HealthBench benchmarks. Earlier, as a Software Development Intern at Dell Technologies, I contributed to deployment orchestration and infrastructure automation using Golang and Python.

