Skip to main content
Maksim MillerMM
Looking for a job

Maksim Miller

@maksmiller

Senior AI/ML engineer building low-latency, cost-efficient LLM inference. Hosted NIM at NVIDIA; before that, search ranking at Yandex Market.

Kazakhstan
Message

What I'm looking for

I'm looking for an early-stage startup where inference is the product. I like problems without an established playbook: new model architectures, odd hardware constraints, workloads that break standard serving assumptions. Remote, contractor.

I'm a senior engineer working on LLM inference infrastructure: making large models serve fast, reliably and cheaply on GPUs.

At NVIDIA I own a serving tier for hosted NIM: 40+ models for 900+ customer teams, about 4.2B tokens a day on ~180 GPUs. The hard part is economics rather than raw throughput. Reworking model placement with scale-to-zero, NVMe weight streaming and multi-LoRA cut GPU-hours per million tokens by 46%. On the latency side, FP8, prefix caching, cache-aware routing and chunked prefill brought p95 time-to-first-token from 920 ms to 410 ms.

I also built the release gate every model, quantization and engine upgrade passes before reaching customers: 1,400 cases scored by an LLM judge calibrated against human labels. I own embedding and reranker serving too, where I cut unsupported RAG claims from 11.4% to 3.9%. Alongside the engineering I plan capacity for the tier, mentor engineers and take on-call.

Before NVIDIA I spent two and a half years on search and ranking at Yandex Market: neural ranking and semantic retrieval, and p99 latency from 210 to 95 ms at 6.8K RPS.

I'm interested in inference optimization, serving economics and evaluation. I'm looking for remote roles on teams where LLM serving is the core product.

Experience

Work history, roles, and key accomplishments

NVIDIA logoNV
Current

Senior Software Engineer, AI Inference

Sep 2023 - Sep 2026 (3 years)

Technical owner of one hosted NIM serving tier: 40+ models, 900+ customer teams, 4.2B tokens daily. Cut GPU-hours per 1M tokens by 46% via scale-to-zero, NVMe weight streaming and multi-LoRA, and TTFT p95 from 920 to 410 ms with FP8, prefix caching and cache-aware routing. Own the release gate for every model and engine upgrade (1,400 cases, calibrated LLM judge) and embedding/reranker serving.

Yandex Market logoYM

Senior Backend Engineer

Yandex Market

Jan 2023 - Aug 2023 (7 months)

Technical owner of the product-search read path at 6.8K RPS peak against a 99.95% SLO. Cut p99 latency from 210 ms to 95 ms, led capacity planning and change reviews, and mentored two junior engineers.

Yandex Market logoYM

Backend Engineer

Yandex Market

Mar 2021 - Dec 2022 (1 year 9 months)

Replaced gradient-boosted ranking with a neural ranker served on ONNX Runtime in the hot path, and added ANN-based semantic retrieval to a lexical-only engine, cutting zero-result queries from 7.4% to 2.1%.

Education

Degrees, certifications, and relevant coursework

ITMO University logoIU

ITMO University

Bachelor of Science, Applied Mathematics and Computer Science

Bachelor of Science in Applied Mathematics and Computer Science from ITMO University.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan