At xAI, I evaluated large language model responses through human assessment and model comparisons, using preference ratings and detailed written feedback to identify strengths and weaknesses.
I tested audio models on voice-command understanding and assessed speech recognition and transcription across dialects, accents, and speaking styles.
I identified dialects in datasets and checked linguistic accuracy and cultural alignment for AI training. I followed detailed evaluation guidelines to produce reliable annotation data and report errors that could support model improvements.
My research project was a systematic review of predictive models for treatment toxicity and survival in hematologic malignancies, with a focus on machine learning for frailty stratification.

