Touchstone

Agentic

GOAL
Can the model drive a multi-step tool loop to a goal.
WHAT IT TESTS
Plan -> act -> observe loops: tool selection, recovery, finishing decisively.
HOW IT GRADES
Agent runs against scripted tool environments, graded on task completion.

Standings — gauntlet benches solved

ModelSolved
claude-sonnet-4.6cloud6/6
qwen3.6-35b-a3b6/6
qwen3.6-35b-a3b-or6/6
qwen3.6-35b-bf166/6
qwen3-coder-30b6/6
qwen3-next-80b6/6
devstral-small-2-25126/6
gemma-4-26b-a4b-it6/6
gpt-oss-20b6/6
granite-4.1-30b6/6
granite-4.1-8b6/6
llama-3.3-70b6/6
mistral-small-3.1-24b-instruct-25036/6
qwen3.5-122b-a10b-cluster6/6
qwen3.5-2b6/6
qwen3.5-4b6/6
qwen3.6-27b6/6
qwen3.5-0.8b5/6
gemini-2.5-flash-litecloud5/6
gpt-oss-120b5/6
phi-4-reasoning-plus2/6
deepseek-r1-distill-qwen-32b1/6

Benches — green = in gauntlet

ladder.agentic.chained_dependencyladder.agentic.conditional_branchladder.agentic.error_recoveryladder.agentic.multi_dependency_planladder.agentic.single_callladder.agentic.stateful_recallgate.agentic.constraint_adherencegate.agentic.diff_validgate.agentic.needle_recallgate.agentic.stop_disciplinegate.agentic.tool_state_multiturnladder.agentic.multistepladder.agentic.ordered_procedureladder.agentic.tool_chain