I'm looking for an early-stage startup where inference is the product. I like problems without an established playbook: new model architectures, odd hardware constraints, workloads that break standard serving assumptions. Remote, contractor.

Maksim Miller
@maksmiller
Senior AI/ML engineer building low-latency, cost-efficient LLM inference. Hosted NIM at NVIDIA; before that, search ranking at Yandex Market.
What I'm looking for
I'm a senior engineer working on LLM inference infrastructure: making large models serve fast, reliably and cheaply on GPUs.
At NVIDIA I own a serving tier for hosted NIM: 40+ models for 900+ customer teams, about 4.2B tokens a day on ~180 GPUs. The hard part is economics rather than raw throughput. Reworking model placement with scale-to-zero, NVMe weight streaming and multi-LoRA cut GPU-hours per million tokens by 46%. On the latency side, FP8, prefix caching, cache-aware routing and chunked prefill brought p95 time-to-first-token from 920 ms to 410 ms.
I also built the release gate every model, quantization and engine upgrade passes before reaching customers: 1,400 cases scored by an LLM judge calibrated against human labels. I own embedding and reranker serving too, where I cut unsupported RAG claims from 11.4% to 3.9%. Alongside the engineering I plan capacity for the tier, mentor engineers and take on-call.
Before NVIDIA I spent two and a half years on search and ranking at Yandex Market: neural ranking and semantic retrieval, and p99 latency from 210 to 95 ms at 6.8K RPS.
I'm interested in inference optimization, serving economics and evaluation. I'm looking for remote roles on teams where LLM serving is the core product.
Experience
Work history, roles, and key accomplishments
Technical owner of one hosted NIM serving tier: 40+ models, 900+ customer teams, 4.2B tokens daily. Cut GPU-hours per 1M tokens by 46% via scale-to-zero, NVMe weight streaming and multi-LoRA, and TTFT p95 from 920 to 410 ms with FP8, prefix caching and cache-aware routing. Own the release gate for every model and engine upgrade (1,400 cases, calibrated LLM judge) and embedding/reranker serving.
Senior Backend Engineer
Yandex Market
Jan 2023 - Aug 2023 (7 months)
Technical owner of the product-search read path at 6.8K RPS peak against a 99.95% SLO. Cut p99 latency from 210 ms to 95 ms, led capacity planning and change reviews, and mentored two junior engineers.
Backend Engineer
Yandex Market
Mar 2021 - Dec 2022 (1 year 9 months)
Replaced gradient-boosted ranking with a neural ranker served on ONNX Runtime in the hot path, and added ANN-based semantic retrieval to a lexical-only engine, cutting zero-result queries from 7.4% to 2.1%.
Education
Degrees, certifications, and relevant coursework
ITMO University
Bachelor of Science, Applied Mathematics and Computer Science
Bachelor of Science in Applied Mathematics and Computer Science from ITMO University.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Salary expectations
Job categories
Interested in hiring Maksim?
You can contact Maksim and 90k+ other talented remote workers on Himalayas.
Message MaksimGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
