Namgyu Youn
@namgyuyoun
Software engineer focused on LLM inference and quantization, delivering memory-efficient serving and optimized inference performance.
What I'm looking for
I’m a software engineer focused on LLM inference and quantization, with experience in memory-efficient serving and inference performance optimization. I’ve identified and removed redundancy in Qwen3-14B ONNX export graphs, then built an INT8 TensorRT deployment pipeline using QDQ calibration to achieve measurable latency and memory gains.
I contribute to open-source across vLLM quantization dispatch, KV cache scheduling, and runtime modules, and have 25+ merged PRs to torchao (maintained by PyTorch). I’ve also upstreamed quantization methods and benchmark infrastructure, designed low-bit tensor types for memory-efficient LLM/VLM inference, and implemented low-bit MatMul kernels in Triton for practical throughput improvements.
Experience
Work history, roles, and key accomplishments
Contributing across quantization dispatch (NVFP4/RoPE), KV cache scheduling, and runtime modules in vLLM.
Deep Learning Engineer
CLIKA
Feb 2026 - Mar 2026 (1 month)
Identified redundancy in a Qwen3-14B ONNX export graph and applied graph fusion (Attention, LayerNorm) to reduce node count by 21%. Built an INT8 TensorRT deployment pipeline via QDQ calibration, achieving 36% lower latency and 41% lower memory vs FP16.
Efficient AI Inference Lab
AerLabs
Sep 2025 - Jan 2026 (4 months)
Presented LLM inference serving research including AWQ (MLSys’24), QServe (MLSys’25), and Flash Attention. Discussed kernel-level implementations for efficient inference.
torchao (ao) Contributor
PyTorch Foundation (Meta)
Jul 2025 - Jan 2026 (6 months)
Upstreamed PTQ benchmarking infrastructure (lm-eval, vLLM) and quantization methods (AWQ/SmoothQuant), and built torchao’s first end-to-end GPU memory benchmark. Designed INT8/MXFP8 tensor subclasses for memory-efficient LLM/VLM inference and implemented low-bit MatMul kernels in Triton (INT4/INT8/MXFP8).
Optimizers
Meta Research
Mar 2025 - May 2025 (2 months)
Improved QAT accuracy by implementing a Hadamard rotation kernel (C++) to reduce quantization error. Profiled distributed Shampoo using PyTorch Profiler and identified root inverse and eigenvector computations as major overhead.
Undergraduate Research Assistant
Sogang University
Nov 2022 - Jun 2023 (7 months)
Built a benchmarking pipeline for YOLO- and DETR-based biological object detection models, evaluating mAP, latency, and throughput to guide model selection. Exported the selected DETR model to ONNX and validated inference correctness for deployment.
Education
Degrees, certifications, and relevant coursework
Naver Connect
Naver Boostcamp AI Tech 7th, Computer Vision
2024 - 2025
Activities and societies: Automated data ingestion and recommendation workflows using Airflow and n8n to enable real-time trend analysis.
Completed Naver Boostcamp AI Tech 7th program focused on practical computer vision and MLOps team projects.
Sogang University
Bachelor of Science, Mathematics
2020 - 2024
Grade: GPA: 3.5/4.3
Earned a B.S. in Mathematics at Sogang University.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Interested in hiring Namgyu?
You can contact Namgyu and 90k+ other talented remote workers on Himalayas.
Message NamgyuGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
