At IIT Mandi, I investigated adversarial answer leakage across LLMs and human tutors. I analyzed 7,800 responses and developed a validated detector that improved agreement from κ = −0.28 to κ = 0.93.
For AMORE, I developed multi-objective reinforcement learning for Phi-3-mini, using LoRA and REINFORCE with adaptive reward weighting across five pedagogical objectives. The approach prevented learning-success collapse, raising it from 4% to 53%, and achieved 70.09% Pedagogy Following.
On my work evaluating lightweight VLMs for accessible video understanding, I studied video descriptions for blind and low-vision users using custom accessibility metrics and two evaluation frameworks. I also evaluated prompting strategies and validated FP32 and INT8 deployment on smartphones.
My research spans LLM alignment, reward modeling, and behavioral evaluation, alongside work on multimodal AI and edge deployment. I have co-authored publications including work presented at ICANN 2026 and IJCNLP-AACL 2025, and an Indian patent application for an assistive navigation and health monitoring system.

