At FlyRank AI, I built a page-prioritization scoring model using production search data, achieving 1.00 precision@10 against a 0.20 weighted-heuristic baseline. I designed leakage-safe features and client-level validation to test the model fairly.
On Ask Your Doc, I built a multi-agent RAG pipeline and benchmarked its configurations through repeated runs. Replication showed that reranking improved quality, while query expansion had the largest latency effect; I recommended a smaller configuration.
I also found a parser bug in my custom grounding judge and excluded it from model selection. Comparing its scores with RAGAS helped me identify that the judge was not measuring faithfulness reliably.
In my model architecture work, I implemented a Transformer, GPT-2, Mixture-of-Experts module, and LSTM, and trained the Transformer across multiple GPUs with PyTorch DDP. I’ve also studied distributed training and LLM inference, including tensor parallelism, ZeRO/FSDP, KV caching, and prefill/decode workloads.

