At Aiverbalyze Technologies Pvt. Ltd., I work with multimodal and speech AI models, including Qwen2.5-Omni and OmniVoice. I contributed to adapting Qwen2.5-Omni and expanding its vocabulary, and explored optimization and pruning for Gemma-based models.
My PhD research at BIT Mesra focused on attention-based visual question answering. I developed models for vision-language understanding, trained and evaluated multimodal models on HPC systems, and optimized lightweight VQA models for the NVIDIA Jetson Orin Nano.
At Comofi Medtech, I implemented a Python pose-estimation algorithm for medical imaging and worked with CT, DRR, C-arm, and multi-camera data. My earlier research included developing a lightweight gait-classification model and deploying it on an embedded device.

