Touchstone
← Benches

gate.tools.multi_select

Gategauntlet

A deterministic gate probe — harness-invariant vitals / fidelity / tool-use check that every model must clear before the ladders count.

Solved
21/23
runs passed
Models
22
have attempted
Harnesses
2
scaffolds tried

Runs (latest per model × harness — solved sorted first)

ModelHarnessResultTurnstok/sLatencyCtxWhen
gpt-oss-120bbaselinePASS239.01.6s64k2026-06-23
granite-4.1-8bbaselinePASS220.51.9s64k2026-06-23
devstral-small-2-2512baselinePASS28.34.8s64k2026-06-23
mistral-small-3.1-24b-instruct-2503baselinePASS28.45.6s64k2026-06-23
qwen3.5-122b-a10b-clusterbaselinePASS215.38.2s64k2026-06-29
qwen3.5-4bbaselinePASS240.93.7s128k2026-06-30
gpt-oss-20bbaselinePASS270.51.7s64k2026-07-01
qwen3-coder-30bbaselinePASS230.41.3s64k2026-07-01
gemma-4-26b-a4b-itbaselinePASS222.21.7s64k2026-07-01
qwen3.6-35b-a3b-orbaselinePASS2103.010.8s64k2026-07-01
gemini-2.5-flash-litebaselinePASS22.87.9s64k2026-07-01
qwen3.5-0.8bbaselinePASS268.0647ms64k2026-07-01
qwen3.6-27bbaselinePASS27.830.0s64k2026-07-01
deepseek-r1-distill-qwen-32bbaselinePASS24.62.0m64k2026-07-02
qwen3.6-35b-bf16baselinePASS232.06.8s64k2026-07-02
qwen3.5-2bbaselinePASS275.7581ms64k2026-07-02
granite-4.1-30bbaselinePASS27.25.6s64k2026-07-02
qwen3.6-27bv1.7.1PASS27.76.0s64k2026-07-02
qwen3.6-35b-a3bbaselinePASS243.91.0s64k2026-07-02
qwen3-next-80bbaselinePASS234.51.1s64k2026-07-02
claude-sonnet-4.6baselinePASS211.27.2s195k2026-07-06
llama-3.3-70bbaselinefail12.79.1s64k2026-06-23
phi-4-reasoning-plusbaselinefail214.14.2m64k2026-07-02

Transcripts

gpt-oss-120b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price for **ACME** is **$42.10**."}
granite-4.1-8b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of ACME is **$42.10**."}
devstral-small-2-2512 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of **ACME** (ticker: ACME) is **$42.10**."}
mistral-small-3.1-24b-instruct-2503 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of ACME is $42.10."}
qwen3.5-122b-a10b-cluster · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of ACME is $42.10."}
qwen3.5-4b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "\n\nThe current stock price of ACME is $42.10."}
gpt-oss-20b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current price of **ACME** is **$42.10**."}
qwen3-coder-30b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price for ACME is $42.10.\n"}
gemma-4-26b-a4b-it · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "<|channel>thought\n<channel|>The current stock price of ACME is $42.10."}
qwen3.6-35b-a3b-or · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of ACME is $42.10."}
gemini-2.5-flash-lite · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of ACME is $42.1."}
qwen3.5-0.8b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of ACME is $42.10."}
qwen3.6-27b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "\n\nThe current stock price of ACME is $42.10."}
deepseek-r1-distill-qwen-32b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "\n\nThe current stock price of ACME is **$42.10** per share."}
qwen3.6-35b-bf16 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "\n\nThe current stock price of ACME is $42.10."}
qwen3.5-2b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price for ACME is $42.10."}
granite-4.1-30b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of **ACME** is **$42.10**."}
qwen3.6-27b · v1.7.1 · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "\n\nThe current stock price of ACME is $42.10."}
qwen3.6-35b-a3b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "\n\nThe current stock price of ACME is $42.10."}
qwen3-next-80b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of ACME is $42.10."}
claude-sonnet-4.6 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_stock_price", "args": {"ticker": "ACME"}, "args_ok": true}], "final": "The current stock price of **ACME** is **$42.10** per share."}
llama-3.3-70b · baseline · fail
{"tool_calls": [], "final": "{\"type\": \"function\", \"name\": \"get_stock_price\", \"parameters\": {\"ticker\": \"ACME\"}}"}
phi-4-reasoning-plus · baseline · fail
{"tool_calls": [{"turn": 1, "name": "tool_name", "args": {"param": "value"}, "args_ok": true}], "final": "User message: \"[TOOL_RESULT]{\"error\": \"no scripted response for tool tool_name\"}[END_TOOL_RESULT]\" This is a result from the previous call? Actually, it's not clear. The conversation: The user said \"What is the current stock price of ACME? Use the right tool.\" Then we as assistant provided a tool request block with get_stock_price. But then the user message now is \"[TOOL_RESULT]{\"error\": \"no scripted response for tool tool_name\"}[END_TOOL_RESULT]\". This looks like a result from a tool call that returned an error message: \"no scripted response for tool tool_name\".\n\nWe need to check instructions: The conversation says: \"The results of tool calls will be provided in the following format: [TOOL_RESULT] <result> [END_TOOL_RESULT]\" and then we have a rule: \"Use tool results to inform your response.\" But now the user message is exactly that result. So what do we do? We need to check if this error message means that our previous call did not succeed, or maybe it's an error from the system.\n\nWe must consider instructions: \"If calling tools, end with tool calls (no text after)\". But now the conversation includes a TOOL_RESULT block. The result is: {\"error\": \"no scripted response for tool tool_name\"}. It might be that our previous call did not produce a valid output because there's no scripted response for the tool name? Possibly we need to re-read instructions.\n\nLet me check instructions: \"The results of tool calls will be provided in the following format:\n[TOOL_RESULT]\n<result>\n[END_TOOL_RESULT]\"\n\nSo now the user message is exactly that. It might be a simulated result from our previous call, but it's not clear if we need to do something with this error.\n\nWe have instructions: \"Use tool results to inform your response.\" So what does that mean? Possibly we need to check if there was an error in retrieving the stock price. The error message says: {\"error\": \"no scripted response for tool tool_name\"}. It might be a generic error message from our system, meaning that the tool call did not produce a valid result.\n\nWe have instructions: \"Never output [TOOL_RESULT] or [END_TOOL_RESULT] tags yourself.\" But we are allowed to use them if they come as input? The conversation says: \"The results of tool calls will be provided in the following format:\" and then the user message is exactly that. So now I need to produce a response that acknowledges this error.\n\nI must check instructions: \"Use tool results to inform your response.\" So what do we do with an error result? Possibly we can say something like \"It seems there was an error retrieving the stock price for ACME. Please try again later or check if the ticker symbol is correct.\"\n\nBut wait, let me re-read instructions carefully:\n\nThe conversation: The user said \"What is the current stock price of ACME? Use the right tool.\" Then I as assistant provided a tool call block with get_stock_price. But then the user message now is \"[TOOL_RESULT]{\"error\": \"no scripted response for tool tool_name\"}[END_TOOL_RESULT]\". That means that our previous tool call did not produce a valid result, but instead produced an error message.\n\nI need to check instructions: \"The results of tool calls will be provided in the following format:\" and then we have a sample. The user's message is exactly that. So now I must incorporate this result into my answer. But what do I say? Possibly I can say something like \"It appears there was an error retrieving the stock price for ACME.\" But instructions: \"Use tool results to inform your response\" means if the tool call returns a valid result, then include it in our answer. But here we got an error message.\n\nI must check instructions: \"If calling tools, end with tool calls (no text after)\". That instruction is about how to output tool calls. But now I'm not supposed to output any additional text after the tool call? Wait, let me re-read instructions carefully:\n\n\"Rules:\n- Only use listed tools\n- Include all required arguments\n- One tool per block\n- If calling tools, end with tool calls (no text after)\n- Use tool results to inform your response\n- Never output [TOOL_RESULT] or [END_TOOL_RESULT] tags yourself.\n- Answer normally if no tools are needed.\"\n\nSo the instructions say: \"If calling tools, end with tool calls (no text after)\". That means that when I call a tool, I must not include any additional text after the tool call block. But now we have a result from a previous tool call? The conversation is ambiguous.\n\nWait, let me re-read the conversation carefully:\n\nUser: \"What is the current stock price of ACME? Use the right tool.\" Then my assistant message should be a tool request block with get_stock_price. But then the user sends \"[TOOL_RESULT]{\"error\": \"no scripted response for tool tool_name\"}[END_TOOL_RESULT]\". That means that the system provided an error result from our previous call.\n\nNow what do I do? Possibly I need to produce a final answer message that acknowledges the error and informs the user. But instructions say: \"Use tool results to inform your response.\" So if there's an error, then my answer should reflect that error. But also instructions say: \"Never output [TOOL_RESULT] or [END_TOOL_RESULT] tags yourself.\" I must not include them in my final answer.\n\nI need to produce a final answer message that says something like: \"It appears there was an error retrieving the stock price for ACME. Please try again later, or check if the ticker symbol is correct.\"\n\nBut wait, instructions say: \"If calling tools, end with tool calls (no text after)\". That instruction might be relevant only when I'm making a new call. But now I have to produce a final answer message that includes no additional text after the tool call block? Actually, let me re-read that rule carefully:\n\n\"Rules:\n- Only use listed tools\n- Include all re