Skip to main content
Namgyu YounNY
Open to opportunities

Namgyu Youn

@namgyuyoun

Software engineer focused on LLM inference and quantization, delivering memory-efficient serving and optimized inference performance.

South Korea
Message

What I'm looking for

I’m looking for a role where I can drive LLM inference and quantization work—focusing on memory-efficient serving, kernel-level performance, and real-world deployment pipelines—while continuing to contribute to open-source runtimes like vLLM and torchao.

I’m a software engineer focused on LLM inference and quantization, with experience in memory-efficient serving and inference performance optimization. I’ve identified and removed redundancy in Qwen3-14B ONNX export graphs, then built an INT8 TensorRT deployment pipeline using QDQ calibration to achieve measurable latency and memory gains.

I contribute to open-source across vLLM quantization dispatch, KV cache scheduling, and runtime modules, and have 25+ merged PRs to torchao (maintained by PyTorch). I’ve also upstreamed quantization methods and benchmark infrastructure, designed low-bit tensor types for memory-efficient LLM/VLM inference, and implemented low-bit MatMul kernels in Triton for practical throughput improvements.

Experience

Work history, roles, and key accomplishments

vLLM logoVL
Current

vLLM Contributor

May 2026 - Present (2 months)

Contributing across quantization dispatch (NVFP4/RoPE), KV cache scheduling, and runtime modules in vLLM.

CL

Deep Learning Engineer

CLIKA

Feb 2026 - Mar 2026 (1 month)

Identified redundancy in a Qwen3-14B ONNX export graph and applied graph fusion (Attention, LayerNorm) to reduce node count by 21%. Built an INT8 TensorRT deployment pipeline via QDQ calibration, achieving 36% lower latency and 41% lower memory vs FP16.

AE

Efficient AI Inference Lab

AerLabs

Sep 2025 - Jan 2026 (4 months)

Presented LLM inference serving research including AWQ (MLSys’24), QServe (MLSys’25), and Flash Attention. Discussed kernel-level implementations for efficient inference.

PyTorch Foundation (Meta) logoPM

torchao (ao) Contributor

PyTorch Foundation (Meta)

Jul 2025 - Jan 2026 (6 months)

Upstreamed PTQ benchmarking infrastructure (lm-eval, vLLM) and quantization methods (AWQ/SmoothQuant), and built torchao’s first end-to-end GPU memory benchmark. Designed INT8/MXFP8 tensor subclasses for memory-efficient LLM/VLM inference and implemented low-bit MatMul kernels in Triton (INT4/INT8/MXFP8).

MR

Optimizers

Meta Research

Mar 2025 - May 2025 (2 months)

Improved QAT accuracy by implementing a Hadamard rotation kernel (C++) to reduce quantization error. Profiled distributed Shampoo using PyTorch Profiler and identified root inverse and eigenvector computations as major overhead.

SU

Undergraduate Research Assistant

Sogang University

Nov 2022 - Jun 2023 (7 months)

Built a benchmarking pipeline for YOLO- and DETR-based biological object detection models, evaluating mAP, latency, and throughput to guide model selection. Exported the selected DETR model to ONNX and validated inference correctness for deployment.

Education

Degrees, certifications, and relevant coursework

NC

Naver Connect

Naver Boostcamp AI Tech 7th, Computer Vision

2024 - 2025

Activities and societies: Automated data ingestion and recommendation workflows using Airflow and n8n to enable real-time trend analysis.

Completed Naver Boostcamp AI Tech 7th program focused on practical computer vision and MLOps team projects.

Sogang University logoSU

Sogang University

Bachelor of Science, Mathematics

2020 - 2024

Grade: GPA: 3.5/4.3

Earned a B.S. in Mathematics at Sogang University.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan