Yuma Watanabe
@yumawatanabe
AI coding-agent evaluator and red-teamer, turning adversarial tests into actionable reports to harden LLM safety.
What I'm looking for
I’m an AI coding-agent evaluator and red-teamer focused on adversarially testing AI coding agents and large language models. I design complex, multi-step and multi-turn instructions that probe reasoning, code changes, and tool use—especially where failures only surface under real difficulty and complexity.
In my contract role (remote), I directed AI coding agents to edit real, git-cloned repositories to expose weaknesses. I prioritize cross-file and cross-component consistency, identifying cases where edits look correct in isolation but break connections between components, then reproduce and document failure modes clearly so engineering teams can act on them.
I also build my own machine-learning models in Python, and I’ve applied that mindset to quantitative finance. In an independent project, I built a risk-focused swing-trading model using the moomoo/futu API, emphasizing minimizing maximum drawdown and validating with a 10-year backtest including look-ahead/data-leakage checks.
I’m a native Japanese speaker with B2-level English, and I enjoy multilingual adversarial evaluation where communication quality matters. I’m drawn to teams that value rigorous test design, structured technical documentation, and safety improvements grounded in measurable failure-mode evidence.
Experience
Work history, roles, and key accomplishments
AI Coding-Agent Evaluator & Red-Teamer
Vendor for Frontier AI Lab
Designed challenging, multi-step instructions to adversarially test AI coding agents editing git-cloned repositories, probing reasoning, code-change, and tool-use weaknesses. Reproduced and documented failure modes with actionable, cross-file consistency findings for engineering teams.
AI Coding-Agent Evaluator
Frontier AI Lab (Vendor)
Designed challenging multi-step, multi-turn prompts to adversarially test AI coding agents against real (git-cloned) repositories, probing reasoning, code changes, and tool use. Reproduced and documented failure modes with actionable reports focused on cross-file and cross-component consistency gaps.
Education
Degrees, certifications, and relevant coursework
Shibaura Institute of Technology
Bachelor of Science, Computer Science
Activities and societies: Focus on machine learning with quantum computing and optical communication; self-directed study in AI safety/adversarial evaluation and quantitative finance.
Pursuing a B.S. in Computer Science at Shibaura Institute of Technology. Focuses on machine learning with quantum computing and optical communication, with self-directed study in AI safety/adversarial evaluation and quantitative finance.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Interested in hiring Yuma?
You can contact Yuma and 90k+ other talented remote workers on Himalayas.
Message YumaGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
