I build speech-to-speech conversational AI systems that connect ASR, LLMs, and TTS. My work includes a pipeline using Faster-Whisper, instruction-fine-tuned TinyLlama, Kokoro TTS, voice activity detection, and waveform, STFT, and Mel-spectrogram analysis.
I'm pursuing an M.Tech in Signal Processing at NIT Calicut and apply deep learning, signal processing, probability, and optimization to speech, language, and EEG systems. I work with PyTorch, Hugging Face Transformers, LangChain, LangGraph, and retrieval-augmented generation.
I've built preference-alignment pipelines spanning SFT, DPO, and RLVR/GRPO, including a GRPO LoRA adapter trained on a custom reasoning and QA dataset. I also built a sparse Mixture-of-Experts GPT from scratch, reducing validation perplexity from 91 to 44 over 10K iterations within a 4GB VRAM budget.
My EEG emotion-recognition work achieved 99.69% accuracy with GRU plus Attention and 99.06% with LSTM plus Attention, with results submitted to an international conference. I have also developed document QA, computer vision, GAN image inpainting, and LLM fine-tuning benchmarks.
