At YarrTech, I evaluate LLM-powered agents across tool calling, API interactions, and multi-step tasks. I design repeatable scenarios and investigate failures so model and product teams can reproduce them.
I built DriftGuard, an LLM API regression detector that runs frozen scenario suites against new model versions. Its logs link failures to the exact prompt, response, and diff, and a GitHub App can open a draft fix PR.
On Recepta, I architected a multi-tenant conversational AI receptionist backend where the application controls booking state. I added a fingerprint-based confirmation gate and shipped voice-note, image-classification, and menu OCR workflows.
For my final-year project at FAST NUCES, I built an end-to-end computer vision system for surgical instrument defect detection, fine-tuning YOLOv5s and deploying the model through a Flask inference API. I also built a synthetic media detection system with a locally run, replayable investigation workflow.

