At Freethought Labs, I direct research in mechanistic interpretability and representation engineering. My first-author paper, “Shared Truth,” was accepted to the NeurIPS 2026 Main Track and tests whether truth representations transfer across 24 model pairs.
I found that closed-form Procrustes alignment matched trained autoencoder adapters on 19 of 21 pairs. I also investigated how injection strength affects transfer and tested the results against native-ceiling, random-weight, and shuffled-adapter baselines.
At Google, I fine-tuned LLM raters for deepfake and text-to-video detection during my YouTube Trust & Safety internship. At Google Maps Platform, I designed and shipped the experimental Air Quality widget for the Google Maps JavaScript API.
My M.S. thesis examined objective sufficiency modeling, including a triadic LLM data-generation pipeline and BERT-family classifiers. I’m pursuing a PhD in computer science and continue to explore how model representations can support interpretable decisions.

