
Gagandeep Suthar
@gagandeepsuthar
I build rigorous benchmarks and evaluation systems for LLMs and coding agents.
What I'm looking for
I've built evaluation systems for LLMs and coding agents at Xelron AI and Airdawg Labs, validating outputs for correctness, reliability, safety, and task compliance. My work includes RLHF feedback, prompt variants, failure analysis, rubrics, Golden Responses, and document-grounded evaluation.
I built LLM-Bench to benchmark 1,000+ prompts across four models and CodeAgent-Eval with 250+ Docker-isolated coding tasks. I also developed RAG-Eval across 50+ source documents and 500+ questions, improving retrieval quality, citation accuracy, and hallucination measurement.
Experience
Work history, roles, and key accomplishments
AI System Evaluation Intern
Airdawg Labs
May 2026 - Present (4 months)
Validate LLM outputs for correctness, completeness, reliability, and adherence to task requirements. Design prompt variants, test cases, and edge scenarios to expose model limitations and document findings.
Software Engineer Intern
Xelron AI
Feb 2025 - Jun 2025 (4 months)
Built Docker-based evaluation pipelines and sandbox challenges for AI coding agents. Evaluated agent conversations and traces, authored rubrics and Golden Responses, and performed pairwise comparisons across frontier LLMs.
Education
Degrees, certifications, and relevant coursework
Indian Institute of Information Technology, Sonepat
B.Tech, Information Technology
2023 -
Pursuing a B.Tech in Information Technology, expected to graduate in 2027.
Availability
Location
Authorized to work in
Portfolio
github.com/Gagan003Salary expectations
Skills
Interested in hiring Gagandeep?
You can contact Gagandeep and 90k+ other talented remote workers on Himalayas.
Message GagandeepGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
