On CalorAI, I collaborated with the founders on a LangGraph agent and built an outcome-based evaluation suite. In adversarial tests, I cut destructive wrong-tool calls from 64% to 16% and improved safety checks.
At PRAN Foundation, I implemented a 3D U-Net pipeline for electron-microscopy neuron segmentation. A data fix lifted Dice from 0.39 to 0.92, and mixed precision delivered about 78x speedup.
At ScaleAI (Outlier), I red-teamed frontier LLMs with adversarial tool-use tasks and created evaluation rubrics and supervised fine-tuning datasets. At Haltdos, I contributed to a no-code automation engine and developed ML systems that reduced manual intervention and fault-detection time.

