Poojith Devan
@poojithdevan
I build benchmarked LLM inference stacks with Triton kernels, quantization, and vLLM.
What I'm looking for
At OneBit, I built packed ternary Triton inference kernels that increased Qwen3-8B decode from 1.48 to 15.5 tok/s while reducing VRAM from 15.3 to 7.4 GiB on an RTX 5070. I benchmark latency, throughput, VRAM, and quality rigorously, including 100% argmax agreement against reference outputs.
I've also implemented integer-only Q1.15 inference for FPU-less embedded hardware at Trebuchet System, reaching 77.71% CIFAR-10 accuracy on ResNet-20. My projects span vLLM serving benchmarks, quantization tradeoffs, and OpenAI-compatible LLM gateways with routing, semantic caching, failover, and rate limiting.
Experience
Work history, roles, and key accomplishments
Research Intern
OneBit
Apr 2026 - Present (5 months)
Built a Triton kernel for packed ternary weights, improving Qwen3-8B decode speed from 1.48 to 15.5 tok/s and reducing VRAM from 15.3 to 7.4 GiB on a 12 GB RTX 5070. Verified 100% argmax agreement with reference and profiled with Nsight Compute to optimize memory access.
Education
Degrees, certifications, and relevant coursework
SRM University
Master of Computer Applications, Generative AI
Pursuing a Master of Computer Applications in Generative AI, with coursework completed remotely and available for full-time work.
O.P. Jindal Global University
Master of Science, AI & Data Science
Completed a Master of Science in AI & Data Science.
SRM University
Bachelor of Computer Applications, Data Science
Completed a Bachelor of Computer Applications in Data Science.
Availability
Location
Authorized to work in
Portfolio
github.com/poojithdevan4Job categories
Interested in hiring Poojith?
You can contact Poojith and 90k+ other talented remote workers on Himalayas.
Message PoojithGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
