
Hassan Shahzad
@hassanshahzad3
I build and evaluate LLM applications, from agentic RAG systems to fine-tuned Text-to-SQL models.
What I'm looking for
I've evaluated frontier LLM outputs at Turing, including Google Gemini, Anthropic Claude, and DeepSeek, assessing correctness, coding quality, and adherence to evaluation protocols.
I authored reference implementations and public and private test suites, and built pytest-based automated testing with 90%+ line and branch coverage across evaluation pipelines.
I built an agentic RAG system for Hugging Face Transformers and PEFT documentation using Python, LangGraph, ChromaDB, FastAPI, Docker, hybrid retrieval, and GitHub source-code lookup. Its evaluation reached 0.81 Faithfulness and 0.94 Answer Relevancy using RAGAS.
I've also fine-tuned Llama 3.1 8B for Text-to-SQL using QLoRA and PEFT, achieving approximately 80% execution-equivalent accuracy, and built AssignGuard, an NLP-based academic similarity detection system.
Experience
Work history, roles, and key accomplishments
Evaluated outputs from frontier LLMs including Google Gemini, Anthropic Claude, and DeepSeek, assessing correctness, coding quality, and adherence to evaluation protocols. Authored ideal reference implementations and designed public/private test suites to benchmark model-generated code across varied difficulty levels.
Education
Degrees, certifications, and relevant coursework
Institute of Management Sciences
Bachelor of Computer Science, Computer Science
2022 -
Pursuing a Bachelor of Computer Science (BCS) at the Institute of Management Sciences, Peshawar, with an expected graduation in July 2026.
Availability
Location
Authorized to work in
Salary expectations
Skills
Interested in hiring Hassan?
You can contact Hassan and 90k+ other talented remote workers on Himalayas.
Message HassanGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
