At DRDO, Research Centre Imarat (RCI) / DRDL, I improved an evaluation pipeline by replacing a scoring method that under-reported model quality with Normalized Wasserstein Distance. I also benchmarked three detectors across GPUs and traced contradictory dashboard metrics back to their data and evaluation sources.
As an Undergraduate Research Intern at IIT Indore, I exposed a false result in a synthetic-noise control that appeared to show embeddings tracking noise, then shipped the pipeline as a CLI package with 120 automated tests.
In my Evaluation Assurance project, I ranked evaluator verdicts using execution checks to catch incorrect PASS verdicts while sending only a portion of cases to human review. I also built an evaluation harness for EDGAR MCP agent tools, where ambiguous queries were refused correctly in graded runs and accuracy was reported as measured.
My projects include multi-agent decision support with human approval gates, deterministic verification, and an audit log, as well as hybrid RAG that abstains on unsupported answers and attributes responses to sources. I back my training, inference, RAG, and agent pipeline work with tests and written failure analysis.

