
Aman Karki
@amankarki
C++/CUDA engineer working on LLM inference in vLLM and llama.cpp, with two CUDA PRs merged into llama.cpp.
What I'm looking for
In llama.cpp's CUDA backend, I contributed two merged PRs: a POOL_1D CUDA kernel that matched the CPU reference across 216 test cases on T4s, and CUDA support for i16/i32 GGML_OP_DUP that fixed a silent CPU fallback. I also have two open PRs in vLLM for Ada FP8 testing and benchmarking.
I'm working on block-FP8 GEMM support for Ada GPUs in vLLM; the kernel runs correctly on an L4, and I'm tuning its performance. I built verbum.cpp, a C++20/CUDA LLM inference engine whose logits match HuggingFace to 3e-5, and Lattice, an embedded vector database with a hand-written HNSW. I write about this work at amankarki.hashnode.dev.
Experience
Work history, roles, and key accomplishments
Open-source C++/CUDA work on LLM inference. llama.cpp CUDA backend: 2 PRs merged (POOL_1D kernel; i16/i32 DUP on CUDA, fixing a silent CPU fallback). vLLM (working on): block-FP8 GEMM for Ada GPUs (#58241), correct on an L4 and being tuned; open PRs #60171 and #59261. Built verbum.cpp (C++20/CUDA inference engine) and Lattice (HNSW vector DB).
Education
Degrees, certifications, and relevant coursework
University of Petroleum and Energy Studies
Bachelor of Technology, Computer Science and Engineering (Cyber Security and Forensics)
Completed a Bachelor of Technology in Computer Science and Engineering with a specialization in Cyber Security and Forensics in November 2024.
Availability
Location
Authorized to work in
Portfolio
aman-portfolio-rho-six.vercel.appSalary expectations
Interested in hiring Aman?
You can contact Aman and 90k+ other talented remote workers on Himalayas.
Message AmanGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
