At Capital One Applied Research, I lead tabular foundation model and LLM-generated synthetic data work from research through open-source, production-usable releases. I’m second author of PersonaLedger, a 30M-transaction synthetic financial benchmark used to evaluate 14 models for illiquidity classification and identity-theft detection.
I’ve built federated learning, differential privacy, secure aggregation, and privacy-leakage evaluation systems across financial and healthcare data. My work includes Capital One’s open-source Federated Model Aggregation framework, a federated learning patent, and LLM-based generation gates that identify risky candidate generators before data is shared.
I also build practical research infrastructure with LangGraph, AWS, Airflow, EKS, MLflow, and Python, while mentoring ML engineers and setting privacy and interpretability review standards.
