I’m a Computer Science graduate from IIIT Kota specializing in AI engineering, software engineering, LLM evaluation, and AI-agent benchmarking.
I have hands-on experience with Terminal-Bench 2 and frontier-model evaluation, where I create challenging, reproducible benchmark tasks across software engineering, debugging, cybersecurity, scientific computing, data processing, machine learning, games, and system administration.
My benchmark engineering work includes Docker-based environments, oracle/reference implementations, deterministic verifiers, automated test suites, reproducibility validation, failure analysis, and reward-hacking prevention. I enjoy designing tasks that require genuine engineering and reasoning rather than simple pattern matching.
Beyond evaluation, I have experience building AI and software systems using Python, FastAPI, Node.js, React, Docker, Linux/Bash, LangChain, CrewAI, RAG, vector databases, LLM APIs, and SQL/NoSQL databases. I have worked through AI/software engineering internships and freelance projects focused on AI agents, evaluation infrastructure, and production software.
I’m interested in working on frontier AI evaluation, coding-agent benchmarks, LLM/agentic systems, AI infrastructure, backend engineering, and challenging software-engineering problems.

