Антон Агейков
@0005282
I build high-performance computer vision and LLM inference systems for resource-constrained hardware.
What I'm looking for
At Huawei, I develop CPU inference systems for LLMs, including SIMD-vectorized fused C++ transformer attention operators that improved end-to-end throughput by 20%. I've also ported vLLM fused operators while preserving a 5x throughput gain and built vLLM, llama.cpp, Lmcache, and Valkey server infrastructure for AI agent coding workflows.
Previously, I trained and deployed YOLO, UNet, and vision transformer models for real-time vision pipelines at Luggar Design Bureau and Neurolumber. My work spans RKNN quantization, object tracking, shared-memory multiprocessing, service-based architecture, ZeroMQ integration, and GPU/NPU optimization, with outcomes including reducing YOLO inference from 300 ms to 40 ms and achieving 95% mAP.
Experience
Work history, roles, and key accomplishments
Developed fused C++ operators for transformer attention via SIMD on ARM, achieving 20% throughput improvement. Ported operators to vLLM, preserving 5x throughput gain, and set up vLLM and llama.cpp servers for AI agent coding.
Education
Degrees, certifications, and relevant coursework
NSU CompTech Winter School
Winter School, Computer Science
2024 - 2024
Completed a project on Lumber Defect Detection using YOLO.
Novosibirsk State University
Bachelor's Degree, Physics
2019 - 2023
Pursued a Bachelor's degree in Physics with a focus on Continuum Mechanics.
NSU CompTech Winter School
Winter School, Computer Science
2022 - 2022
Completed a project on COVID-19 Incidence Prediction using time series analysis.
Availability
Location
Authorized to work in
Job categories
Interested in hiring Антон?
You can contact Антон and 90k+ other talented remote workers on Himalayas.
Message АнтонGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
