At Huawei R&D, I proposed a DPD architectural layer adopted into a next-generation chip, cutting ASIC area and power by about 22% at the same signal quality. I also initiated LLM-driven neural architecture search, finding architectures that used a further ~20% less hardware.
I built a PyTorch DPD training framework from scratch with custom C++/CUDA kernels, running 5–15× faster than the PyTorch baseline. This brought testing architectures from adjacent RF-ML tasks down from weeks to days.
At Smart Engines, I designed LRFR, a CT-specific neural layer that delivered higher accuracy than prior state of the art at about 10× lower runtime on AAPM Low Dose CT Grand Challenge data. I also contributed to neural reconstruction that made industrial CT scans up to 2× faster through shorter per-frame X-ray exposure.
My research includes a 2026 preprint on combining linear and compressed sparse attention, with faster training and prefill than pure CSA at matched quality. I hold a Ph.D. from MIPT and have published nine peer-reviewed papers.

