Touchstone

Coding

GOAL
Can the model write and fix code that actually runs.
WHAT IT TESTS
Code generation and bug-fixes across focused exercises.
HOW IT GRADES
Generated code executed against fixed test cases; pass = tests green.

Standings — gauntlet benches solved

ModelSolved
gpt-oss-20b24/24
qwen3.6-35b-a3b24/24
gemma-4-26b-a4b-it24/24
qwen3.5-4b16/24
qwen3.6-27b12/12
claude-sonnet-4.6cloud6/6
deepseek-r1-distill-qwen-32b6/6
granite-4.1-30b6/6
llama-3.3-70b6/6
qwen3.6-35b-a3b-or6/6
qwen3-next-80b6/6
devstral-small-2-25125/6
qwen3-coder-30b5/6
qwen3.5-122b-a10b-cluster5/6
qwen3.6-35b-bf165/6
gpt-oss-120b5/6
mistral-small-3.1-24b-instruct-25035/6
gemini-2.5-flash-litecloud4/6
phi-4-reasoning-plus4/6
granite-4.1-8b4/6
qwen3.5-2b2/6
qwen3.5-0.8b0/6

Benches — green = in gauntlet

ladder.coding.expr_evalladder.coding.fix_binary_searchladder.coding.lru_cacheladder.coding.merge_intervalsladder.coding.topological_sortladder.coding.wildcard_match