At Maksasoft, I developed a zero-shot voice cloning pipeline that works from a single short audio sample and powers AI voice replies in English and 11 Indian languages. I also implemented audio quality checks to select cleaner speech segments before cloning.
I built a Python and FastAPI backend integrating an LLM, speech-to-text APIs, and voice cloning, with real-time WebSocket streaming. I deployed voice and speech models as a serverless GPU service on Modal, and my projects include B2B funnel analytics, retail vendor analysis, and a vehicle insurance prediction service.

