At AMD, via TalTech, I work on Vitis AI/ML compiler flows that map machine-learning kernels and operators to NPU hardware on Strix and VEK embedded devices. I optimize model offloading, resource utilization, and execution readiness.
I quantize and lightweight architectures using methods including INT8, FP16, PTQ, and QAT, and analyze vision and transformer models for deployment. I also optimize YOLO-based detectors for latency and FPS under NPU constraints.
I built and maintained evaluation and automation infrastructure spanning 1,500+ models, with ONNX evaluation, memory/MAC analysis, and validity checks. This cut evaluation turnaround by over 80%.
I also build RAG, MCP servers, and plan-execute multi-agent tooling to expose internal ML and NPU tools to LLM agents. Earlier, as a Research Intern at Indian Institute of Foreign Trade, I summarized research on blockchain in supply chains and used AHP and DEMATEL for decision analysis.

