Skip to main content
GS
Looking for a job

Gagandeep Suthar

@gagandeepsuthar

I build rigorous benchmarks and evaluation systems for LLMs and coding agents.

India
Message

What I'm looking for

I'm looking for opportunities to build and improve rigorous LLM and coding-agent evaluation workflows, including benchmarks, testing, failure analysis, and reliable AI systems.

I've built evaluation systems for LLMs and coding agents at Xelron AI and Airdawg Labs, validating outputs for correctness, reliability, safety, and task compliance. My work includes RLHF feedback, prompt variants, failure analysis, rubrics, Golden Responses, and document-grounded evaluation.

I built LLM-Bench to benchmark 1,000+ prompts across four models and CodeAgent-Eval with 250+ Docker-isolated coding tasks. I also developed RAG-Eval across 50+ source documents and 500+ questions, improving retrieval quality, citation accuracy, and hallucination measurement.

Experience

Work history, roles, and key accomplishments

AL
Current

AI System Evaluation Intern

Airdawg Labs

May 2026 - Present (4 months)

Validate LLM outputs for correctness, completeness, reliability, and adherence to task requirements. Design prompt variants, test cases, and edge scenarios to expose model limitations and document findings.

Education

Degrees, certifications, and relevant coursework

Indian Institute of Information Technology, Sonepat logoIS

Indian Institute of Information Technology, Sonepat

B.Tech, Information Technology

2023 -

Pursuing a B.Tech in Information Technology, expected to graduate in 2027.

Tech stack

Software and tools used professionally

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan