At 10xEngineers, I engineered a CUDA Image Signal Processor pipeline for NVIDIA GPUs and Jetson edge platforms. It improved image quality by 10% and downstream AI model accuracy by 8% while sustaining 4K 60 FPS.
I also designed Quantex, an automated Python and PyTorch quantization framework. It reduced LLM size by up to 20% and Stable Diffusion model size by 25%, within a 1% accuracy drop margin.
Earlier at 10xEngineers, I built edge deployment pipelines and ported YOLOv7 and Moondream to Mojo, achieving inference speedups of 13% and 5%. I also applied structured pruning, knowledge distillation, and quantization to deploy low-latency models on resource-constrained hardware.
On CuQwen, my open-source project, I engineered a C++/CUDA inference engine for low-latency Qwen2.5 execution and developed custom FP16 kernels. My work also includes LLVM code-analysis research at LUMS and an upstream LLDB contribution that added colorized symbol-search output.

