I build machine learning systems for clinical and research data, from LLM-driven agent pipelines to deep learning models. At DataThink, I lead a three-person team automating raw clinical data into CDISC-compliant datasets with confidence-scored, citation-backed decisions for human review.
I maintain R packages that expose more than 20 statistical methods as tools for an LLM agent, serving as the statistical engine across client work. I also built a HIPAA/GDPR de-identification pipeline using spaCy NER and regex to process roughly 20,000 patient messages with zero PHI leakage in review.
Previously, I fine-tuned a ResNet-50 segmentation model with Kitware and trained a PyTorch CNN on a 2.4 million-example genomics benchmark at Duke, achieving about 96% in-distribution accuracy while detecting out-of-distribution inputs.
