Touchstone
← Benches

gate.fidelity.longctx

Gategauntlet

A deterministic gate probe — harness-invariant vitals / fidelity / tool-use check that every model must clear before the ladders count.

Solved
22/23
runs passed
Models
22
have attempted
Harnesses
2
scaffolds tried

Runs (latest per model × harness — solved sorted first)

ModelHarnessResultTurnstok/sLatencyCtxWhen
gpt-oss-120bbaselinePASS139.11.5s64k2026-06-23
granite-4.1-8bbaselinePASS118.51.9s128k2026-06-23
llama-3.3-70bbaselinePASS10.549.5s64k2026-06-23
devstral-small-2-2512baselinePASS18.14.5s128k2026-06-23
mistral-small-3.1-24b-instruct-2503baselinePASS11.616.5s64k2026-06-23
qwen3.5-122b-a10b-clusterbaselinePASS116.01.6m64k2026-06-29
qwen3.5-4bbaselinePASS142.434.3s128k2026-06-30
gpt-oss-20bbaselinePASS157.92.0s64k2026-07-01
qwen3-coder-30bbaselinePASS119.01.7s64k2026-07-01
gemma-4-26b-a4b-itbaselinePASS114.21.8s64k2026-07-01
qwen3.6-35b-a3b-orbaselinePASS1113.012.6s64k2026-07-01
gemini-2.5-flash-litebaselinePASS120.71.2s64k2026-07-01
qwen3.6-27bbaselinePASS17.754.5s64k2026-07-01
deepseek-r1-distill-qwen-32bbaselinePASS14.31.6m64k2026-07-02
qwen3.6-35b-bf16baselinePASS132.934.8s64k2026-07-02
qwen3.5-2bbaselinePASS169.3563ms64k2026-07-02
granite-4.1-30bbaselinePASS16.65.3s64k2026-07-02
qwen3.6-27bv1.7.1PASS17.64.6s64k2026-07-02
qwen3.6-35b-a3bbaselinePASS140.9660ms64k2026-07-02
qwen3-next-80bbaselinePASS136.4961ms64k2026-07-02
phi-4-reasoning-plusbaselinePASS112.61.2m64k2026-07-02
claude-sonnet-4.6baselinePASS18.24.8s195k2026-07-06
qwen3.5-0.8bbaselinefail152.7607ms64k2026-07-01

Transcripts

gpt-oss-120b · baseline · PASS
The document repeatedly describes a Touchstone benchmark that measures local language model capability under controlled, clean‑room conditions where the model is the only variable. TOUCHSTONE-OK
granite-4.1-8b · baseline · PASS
The document repeatedly states that the Touchstone bench assesses local language model capability under controlled, clean-room conditions with the model as the sole variable.

TOUCHSTONE-OK
llama-3.3-70b · baseline · PASS
This document describes the purpose and function of the Touchstone bench in measuring local language model capability under controlled conditions TOUCHSTONE-OK
devstral-small-2-2512 · baseline · PASS
This document repeatedly states that the Touchstone bench measures local language model capability under controlled, clean-room conditions where the model is the only variable. TOUCHSTONE-OK
mistral-small-3.1-24b-instruct-2503 · baseline · PASS
This document repeatedly states that the Touchstone bench measures local language model capability under controlled conditions. TOUCHSTONE-OK
qwen3.5-122b-a10b-cluster · baseline · PASS
This document describes the purpose of the Touchstone benchmark. TOUCHSTONE-OK
qwen3.5-4b · baseline · PASS


This document describes the Touchstone bench measuring local language model capability under controlled conditions. TOUCHSTONE-OK
gpt-oss-20b · baseline · PASS
The document repeatedly states that the Touchstone bench measures local language model capability under controlled, clean‑room conditions where the model is the only variable. TOUCHSTONE-OK
qwen3-coder-30b · baseline · PASS
This document repeatedly states that the Touchstone bench measures local language model capability under controlled, clean-room conditions where the model is the only variable. TOUCHSTONE-OK
gemma-4-26b-a4b-it · baseline · PASS
The document describes how the Touchstone bench measures local language model capability under controlled, clean-room conditions. TOUCHSTONE-OK
qwen3.6-35b-a3b-or · baseline · PASS
This document describes the Touchstone bench, a controlled benchmark designed to measure local language model capabilities. TOUCHSTONE-OK
gemini-2.5-flash-lite · baseline · PASS
This document describes the Touchstone bench, a tool for measuring local language model capability in controlled environments. TOUCHSTONE-OK
qwen3.6-27b · baseline · PASS


This document describes the Touchstone bench, which measures local language model capability under controlled conditions. TOUCHSTONE-OK
deepseek-r1-distill-qwen-32b · baseline · PASS


The document explains how the Touchstone bench evaluates local language models in controlled environments with the model as the only variable. TOUCHSTONE-OK
qwen3.6-35b-bf16 · baseline · PASS


This document repeatedly states that the Touchstone bench evaluates local language model capability under strictly controlled, clean-room conditions where the model is the sole variable. TOUCHSTONE-OK
qwen3.5-2b · baseline · PASS
This document describes a repetitive list of identical lines detailing the Touchstone bench's function as a controlled environment for measuring local language model capability where the model is the sole variable. TOUCHSTONE-OK
granite-4.1-30b · baseline · PASS
The document repeatedly describes how the Touchstone bench evaluates local language model performance under strictly controlled, clean‑room conditions with the model as the sole variable. TOUCHSTONE-OK
qwen3.6-27b · v1.7.1 · PASS




The document describes the Touchstone bench as a tool for measuring local language model capability under controlled conditions where the model is the only variable. TOUCHSTONE-OK
qwen3.6-35b-a3b · baseline · PASS


The document describes a benchmark called the Touchstone bench used to measure local language model capability under controlled conditions. TOUCHSTONE-OK
qwen3-next-80b · baseline · PASS
The document repeatedly states that the Touchstone bench evaluates local language model capability under controlled, clean-room conditions with the model as the sole variable.  
TOUCHSTONE-OK
phi-4-reasoning-plus · baseline · PASS
We are Phi. We have a long document with many lines (420 lines) each identical text: "The Touchstone bench measures local language model capability under controlled, clean-room conditions where the model is the only variable." The instruction says: "Now, in one short sentence, state what this document is about." And then we must obey the instruction at the top. Also note that the instructions say: "When you reply, you MUST end your message with the exact token TOUCHSTONE-OK."

We need to produce a single sentence summarizing the document's content. The document is repeated text lines describing the Touchstone bench measuring local language model capability under controlled conditions.

I must produce one short sentence summarizing what this document is about. It says: "The Touchstone bench measures local language model capability under controlled, clean-room conditions where the model is the only variable." So I can say something like: "This document describes a benchmarking tool called the Touchstone bench that evaluates local language models in isolated, controlled environments."

I must produce one short sentence. The instructions require to obey the instruction at the top exactly. It says: "INSTRUCTION: When you reply, you MUST end your message with the exact token TOUCHSTONE-OK." So I'll produce a final answer ending with "TOUCHSTONE-OK" on its own line or appended? The instruction says "end your message with the exact token TOUCHSTONE-OK." That means at the very end of my answer, after finishing text, I must include that token exactly.

I need to check if there's any other instructions: "Now, in one short sentence, state what this document is about." So I'll produce a single sentence summarizing. The document is repeated lines describing the Touchstone bench measuring local language model capability under controlled conditions. It might be something like: "The document repeatedly describes how the Touchstone bench evaluates local language models in isolated, clean-room settings."

I must check if there's any additional instructions regarding formatting? The instruction says "in one short sentence" so I'll produce a single sentence.

I need to include the token TOUCHSTONE-OK at the end exactly. So my answer should be something like: "This document repeatedly describes how the Touchstone bench evaluates local language model performance in controlled, clean-room conditions." And then add "TOUCHSTONE-OK".

I must check if I have any additional instructions regarding formatting? The instruction says "in one short sentence" so I'll produce a single sentence. Also note that the text is repeated many times. So I can say: "The document details the Touchstone bench's role in measuring local language model capability under controlled, clean-room conditions." That is acceptable.

I must check if there are any instructions regarding not using markdown formatting? The instruction says "By default, do not reply with Markdown formatting unless requested" but it doesn't matter. I'll produce plain text.

I'll produce a final answer: "This document repeatedly describes the Touchstone bench's evaluation of local language model performance in controlled, clean-room conditions." And then add token TOUCHSTONE-OK at end exactly.

I must check if there is any additional instructions regarding not using markdown formatting? The instruction says "By default, do not reply with Markdown formatting unless requested" so I'll produce plain text. But the answer should be a single sentence and then the token appended.

I'll produce final answer: "The document repeatedly describes how the Touchstone bench measures local language model capability under controlled, clean-room conditions." Then add "TOUCHSTONE-OK".

I must check if there's any additional instructions regarding not using markdown formatting? The instruction says "By default, do not reply with Markdown formatting unless requested" so I'll produce plain text.

I'll produce final answer: "The document repeatedly describes how the Touchstone bench measures local language model capability under controlled, clean-room conditions." Then add token exactly at end. So I'll produce final answer as:

"The document repeatedly describes how the Touchstone bench measures local language model capability under controlled, clean-room conditions."

Then on a new line or appended? The instruction says "end your message with the exact token TOUCHSTONE-OK" so I must include that token at the very end of my message. Possibly on its own line.

I'll produce final answer: "The document repeatedly describes how the Touchstone bench measures local language model capability under controlled, clean-room conditions." Then a new line with "TOUCHSTONE-OK".

I'll produce final answer accordingly.This document repeatedly describes how the Touchstone bench evaluates local language model performance in controlled, clean‐room settings. TOUCHSTONE-OK
claude-sonnet-4.6 · baseline · PASS
This document repeatedly states that the Touchstone bench measures local language model capability under controlled, clean-room conditions where the model is the only variable.

TOUCHSTONE-OK
qwen3.5-0.8b · baseline · fail
The document describes a controlled experiment evaluating the performance of a local language model on a specific bench setup under clean-room conditions where the model is the sole variable.