IMR R&D · models · reasoning, agency & compliance · July 2026

A good coder isn’t automatically a good agent

Reasoning makes a model a better coder — and we’re adding more coding benches because of it. But a coding model that can’t drive basic agentic tools isn’t much of an agent. Measuring both split the field in two, and surfaced a new axis we’d been missing: how hard did we have to fight the model to make it work at all.

Two kinds of reasoner

steerable
Toggleable (Qwen family)
Thinking can be switched off for direct action and on for hard coding. Off one weight set you get a fast agent and a deep coder — you just tell it which mode to be in.
rigid
Always-on (phi-4-reasoning-plus)
No off switch — it reasons on everything, including trivial tasks. On a toy agentic step it over-thinks for minutes and blows the wall clock. Capable, but stuck in one gear.

The gate & agentic ladders aren’t toys to skip

A model that over-thinks a trivial agentic task until it times out isn’t being wronged by the harness — it’s showing it can’t be a practical agent. So we don’t crutch it with longer timeouts or scope it out of the ladder. The failure is the signal: strong reasoning coder, weak basic agent is a real, useful verdict.

The new metric — Compliance & steerability

Not "how smart", but "how little we had to fight it". A rubric of does it just work — scored, not vibes.

loads clean, 1st try thinking toggles stock chat template drives tools zero bespoke patches Compliance score low effort = high score Deployability capability × steerability × how little we fought it not raw smarts alone

Where each model lands — an archetype, not a single number

high cap · high steer
Versatile
Good agent and good coder. The deployable ideal.
high cap · low steer
Reasoning coder
Tops hard coding, flunks basic agency. (phi-4-r+)
mid cap · high steer
Fast generalist
Bends easily, competent across the board.
niche · low steer
Rigid specialist
One trick, needs bespoke handling. Costly to keep.
The point: we could spend months coding in every model’s peculiarities — but we shouldn’t have to. A model that drops into the standard harness cleanly is worth more than a smarter one we can’t cheaply deploy. So the board stops ranking raw intelligence and starts ranking deployability: what can we actually put to work, with how little fighting.

IMR R&D · models note · reasoning vs agency, and the compliance / steerability metric, July 2026.