At StellarAI, I designed end-to-end evaluation scenarios for LLMs, including prompts, tool-call selection, realistic databases, and gold-standard trajectories targeted at an 80%+ model failure rate.
I create coding prompts across C++, Python, JavaScript, and Bash, then write multi-tier rubrics to compare outputs from competing models on correctness, completeness, readability, and task alignment. I also build virtual-assistant simulations, annotate cross-OS navigation tasks, review peer submissions, and make final accept/reject decisions.
I bring a Python and software QA background, including a seven-stage astrometry pipeline that achieved 92.7% completeness with 0% false positives on a synthetic 274-source image.
