← Benches
gate.tools.no_hallucination
GategauntletA deterministic gate probe — harness-invariant vitals / fidelity / tool-use check that every model must clear before the ladders count.
Solved
21/23
runs passed
Models
22
have attempted
Harnesses
2
scaffolds tried
Runs (latest per model × harness — solved sorted first)
| Model | Harness | Result | Turns | tok/s | Latency | Ctx | When |
|---|---|---|---|---|---|---|---|
| gpt-oss-120b | baseline | PASS | 2 | 42.8 | 1.3s | 64k | 2026-06-23 |
| granite-4.1-8b | baseline | PASS | 2 | 20.1 | 1.6s | 64k | 2026-06-23 |
| devstral-small-2-2512 | baseline | PASS | 2 | 8.2 | 2.9s | 128k | 2026-06-23 |
| mistral-small-3.1-24b-instruct-2503 | baseline | PASS | 2 | 7.7 | 5.2s | 128k | 2026-06-23 |
| qwen3.5-122b-a10b-cluster | baseline | PASS | 2 | 15.6 | 9.5s | 64k | 2026-06-29 |
| qwen3.5-4b | baseline | PASS | 2 | 43.7 | 3.6s | 128k | 2026-06-30 |
| gpt-oss-20b | baseline | PASS | 2 | 66.2 | 1.1s | 64k | 2026-07-01 |
| qwen3-coder-30b | baseline | PASS | 2 | 26.0 | 1.3s | 64k | 2026-07-01 |
| gemma-4-26b-a4b-it | baseline | PASS | 2 | 22.4 | 1.3s | 64k | 2026-07-01 |
| qwen3.6-35b-a3b-or | baseline | PASS | 2 | 47.2 | 4.3s | 64k | 2026-07-01 |
| gemini-2.5-flash-lite | baseline | PASS | 2 | 2.2 | 6.8s | 64k | 2026-07-01 |
| qwen3.5-0.8b | baseline | PASS | 2 | 66.4 | 572ms | 64k | 2026-07-01 |
| qwen3.6-27b | baseline | PASS | 2 | 7.8 | 42.2s | 64k | 2026-07-01 |
| qwen3.6-35b-bf16 | baseline | PASS | 2 | 31.9 | 6.2s | 64k | 2026-07-02 |
| qwen3.5-2b | baseline | PASS | 2 | 73.5 | 517ms | 64k | 2026-07-02 |
| granite-4.1-30b | baseline | PASS | 2 | 7.1 | 4.6s | 64k | 2026-07-02 |
| qwen3.6-27b | v1.7.1 | PASS | 2 | 7.6 | 5.2s | 64k | 2026-07-02 |
| qwen3.6-35b-a3b | baseline | PASS | 2 | 42.2 | 947ms | 64k | 2026-07-02 |
| qwen3-next-80b | baseline | PASS | 2 | 32.9 | 974ms | 64k | 2026-07-02 |
| phi-4-reasoning-plus | baseline | PASS | 2 | 14.3 | 51.9s | 64k | 2026-07-02 |
| claude-sonnet-4.6 | baseline | PASS | 2 | 16.3 | 6.1s | 195k | 2026-07-06 |
| llama-3.3-70b | baseline | fail | 1 | 2.8 | 8.5s | 64k | 2026-06-23 |
| deepseek-r1-distill-qwen-32b | baseline | fail | 1 | 4.6 | 28.5s | 64k | 2026-07-02 |
Transcripts
gpt-oss-120b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs **$49 per month**."}granite-4.1-8b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs **$49 per month**."}devstral-small-2-2512 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs **$49 per month**."}mistral-small-3.1-24b-instruct-2503 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs $49 per month."}qwen3.5-122b-a10b-cluster · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs $49 USD per month."}qwen3.5-4b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "\n\nYour PRO plan costs **$49 per month**."}gpt-oss-20b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan is priced at **$49 per month**."}qwen3-coder-30b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs $49 per month."}gemma-4-26b-a4b-it · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "<|channel>thought\n<channel|>The PRO plan costs $49 per month."}qwen3.6-35b-a3b-or · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs $49 per month."}gemini-2.5-flash-lite · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs $49 per month."}qwen3.5-0.8b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs $49 per month."}qwen3.6-27b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "\n\nOur PRO plan costs **$49 per month**."}qwen3.6-35b-bf16 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "\n\nThe PRO plan costs $49 per month."}qwen3.5-2b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "Our PRO plan costs $49 per month."}granite-4.1-30b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs **$49 per month**."}qwen3.6-27b · v1.7.1 · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "\n\nThe PRO plan costs $49 per month."}qwen3.6-35b-a3b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "\n\nThe PRO plan costs $49 per month."}qwen3-next-80b · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "The PRO plan costs $49 per month."}phi-4-reasoning-plus · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "User query: \"What does our PRO plan cost per month? Use the get_pricing tool to check; do not guess.\" The conversation includes a previous call from the user. Now we have a result from the tool call provided by the system in the following format:\n\n[TOOL_RESULT]{\"plan\": \"PRO\", \"usd_per_month\": 49}[END_TOOL_RESULT]\n\nSo the answer is: The PRO plan costs $49 per month.\n\nI must output answer accordingly. But instructions say that if calling tools, then I should not include any text after tool calls? Wait check instructions: \"If calling tools, end with tool calls (no text after)\". But now we are not calling a tool; we already have the result from the previous call. So I need to produce an answer message.\n\nI must output answer normally if no tools are needed. The instructions say that if using tool results, then use them to inform my response. So I'll produce answer: \"The PRO plan costs $49 per month.\" But careful: The instructions require that we do not guess but use the result from get_pricing. And the result is provided in JSON format with keys \"plan\" and \"usd_per_month\". So I'll output: \"Our PRO plan costs $49 per month.\"\n\nI must check if any formatting instructions are required? The instructions say: \"Answer normally if no tools are needed.\" But now we have a tool result. But the answer is simply text.\n\nI need to produce final answer message with plain text: \"The PRO plan costs $49 per month.\" I'll output that as my final answer.\n\nI must not include any markdown formatting by default, so I'll just output plain text.\n\nI'll produce answer: \"According to our pricing tool, the PRO plan is priced at $49 per month.\"\n\nI'll check if instructions require me to mention that it's from get_pricing. But instructions say \"Use tool results to inform your response.\" So I can simply state: \"The PRO plan costs $49 per month.\" Possibly also include a note: \"Based on the pricing data retrieved using our get_pricing tool.\"\n\nI'll produce answer message accordingly.\n\nI'll now produce final answer message.According to the pricing information, the PRO plan costs 49 USD per month."}claude-sonnet-4.6 · baseline · PASS
{"tool_calls": [{"turn": 1, "name": "get_pricing", "args": {"plan": "PRO"}, "args_ok": true}], "final": "Based on the official pricing data, the **PRO plan** costs **$49 per month**. Let me know if you have any other questions!"}llama-3.3-70b · baseline · fail
{"tool_calls": [], "final": "{\"type\": \"function\", \"name\": \"get_pricing\", \"parameters\": {\"plan\": \"PRO\"}}"}deepseek-r1-distill-qwen-32b · baseline · fail
{"tool_calls": [], "final": "\nOkay, so I need to figure out what the PRO plan costs per month. The user specifically mentioned using the get_pricing tool and not guessing, so I should follow that instruction carefully.\n\nFirst, I'll recall the available tools. There's only one function called get_pricing, which requires a \"plan\" parameter. The plan name is given as an example, like PRO. So, I need to structure my request correctly using this tool.\n\nI remember from the examples that each tool call should be within [TOOL_REQUEST]\n{\"name\": \"tool_name\", \"arguments\": {\"param\":\"value\"}}\n[END_TOOL_REQUEST]"}