Agentic
GOAL
Can the model drive a multi-step tool loop to a goal.
WHAT IT TESTS
Plan -> act -> observe loops: tool selection, recovery, finishing decisively.
HOW IT GRADES
Agent runs against scripted tool environments, graded on task completion.
Standings — gauntlet benches solved
| Model | Solved | |
|---|---|---|
| claude-sonnet-4.6 | cloud | 6/6 |
| qwen3.6-35b-a3b | 6/6 | |
| qwen3.6-35b-a3b-or | 6/6 | |
| qwen3.6-35b-bf16 | 6/6 | |
| qwen3-coder-30b | 6/6 | |
| qwen3-next-80b | 6/6 | |
| devstral-small-2-2512 | 6/6 | |
| gemma-4-26b-a4b-it | 6/6 | |
| gpt-oss-20b | 6/6 | |
| granite-4.1-30b | 6/6 | |
| granite-4.1-8b | 6/6 | |
| llama-3.3-70b | 6/6 | |
| mistral-small-3.1-24b-instruct-2503 | 6/6 | |
| qwen3.5-122b-a10b-cluster | 6/6 | |
| qwen3.5-2b | 6/6 | |
| qwen3.5-4b | 6/6 | |
| qwen3.6-27b | 6/6 | |
| qwen3.5-0.8b | 5/6 | |
| gemini-2.5-flash-lite | cloud | 5/6 |
| gpt-oss-120b | 5/6 | |
| phi-4-reasoning-plus | 2/6 | |
| deepseek-r1-distill-qwen-32b | 1/6 |
Benches — green = in gauntlet
ladder.agentic.chained_dependencyladder.agentic.conditional_branchladder.agentic.error_recoveryladder.agentic.multi_dependency_planladder.agentic.single_callladder.agentic.stateful_recallgate.agentic.constraint_adherencegate.agentic.diff_validgate.agentic.needle_recallgate.agentic.stop_disciplinegate.agentic.tool_state_multiturnladder.agentic.multistepladder.agentic.ordered_procedureladder.agentic.tool_chain