At ProArch, I architected a reinforcement learning platform for sequential decision-making, training actor-critic agents that improved long-horizon task success by 34% over the legacy heuristic baseline. I also applied RLHF and DPO to fine-tune large language models, improving human-preference win rate by 27% and reducing hallucination rate by 31%.
Previously, at Synack, I built reinforcement learning agents for adaptive security testing that increased unique vulnerability discovery by 28% per engagement. At Meta, I built Python and SQL analytics pipelines and developed early machine learning models for engagement and churn forecasting.

