I've evaluated 500+ LLM prompt-response pairs at Ethara AI, identifying reasoning failures, hallucinations, and alignment inconsistencies while supporting RLHF and post-training datasets.
I also built Research Pilot AI, a FastAPI-based RAG platform with sub-800ms retrieval and automated LLM fallback, and developed a ResNet-50 weed-detection platform that reached 94.7% accuracy.

