Skip to main content
Back to Insights

Clean Exit, Wrong Record

A support agent can close the ticket, return no tool error, and still leave the wrong record. Microsoft and Hugging Face's ThinkingBox (Oct 3, 2026) measured that gap. Gate the customer-visible step on a readback, not the model's "resolved."

An agent can do the careful work, close the ticket, and still be wrong.

On October 3, Microsoft and Hugging Face published ThinkingBox. It grades agents on the records they leave behind, not the sentences they generate. Their walkthrough is an ordinary support case: a delayed kitchen appliance, nine tool calls, the policy read correctly, the ticket closed as resolved, and a polite closer. The required end state was on hold. The customer never got a real answer. A grader watching tool calls would pass it. The database would not.

That is not a leaderboard story. It is the bug already in production agents that treat a clean exit as proof of work.

Receipt Gate says the last write has to leave a skim-worthy receipt before the next one fires. ThinkingBox is the colder half of the same rule. The receipt has to match the record. The model's last line does not get a vote.

Quick use case

Situation. A support team shipped a delivery-exception agent. A stuck shipment comes in. The agent may read the order, check the carrier, open or update a ticket, and reply to the customer. The eval was a golden transcript: the right tools fired, the reply was polite, and the last line said the case was resolved. Staging looked done.

What broke. Put that published retail task on their stack. Tools return success. The ticket lands in solved. The customer gets the closer. The carrier exception is still open, so the required end state is hold, not solved. Nothing in the eval looks at that field. The run is green because the trajectory is tidy. The record is not.

The ThinkingBox writeup is blunt about how often that tidy exit is a costume. In a common-set ablation — 121,680 valid trials across 12 models — 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. The checks still found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those findings overlap. A green trace is not a passed state.

The fix.

  1. Name the end state in fields — status, required effects, forbidden extras. Not "the reply sounds resolved."
  2. Read the record back after the last mutating tool, from the system of record, not from the model's summary.
  3. Latch the customer-visible step on that readback. No send until the fields match.
  4. Receipt the diff — what the record shows, what it was supposed to show, pass or fail in one skim.
  5. Repeat the golden task from a clean start until a one-off pass stops counting as a ship.

Same tools. Same model. Different definition of done. They have not measured the lift from checking terminal state before you commit. You still have to run it on your own records.

Clean Exit Stack: Five layers showing Trajectory, Model line, Readback with teal highlight, Latch with amber block, and Receipt with field diff - the system for gating customer-visible actions on database state verification

A trajectory is a claim

The line worth keeping: a trajectory is a claim. Database state is the evidence. Repetition is the trust test.

Pass@1 says how a model usually does. "Solved at least once in 20 tries" says whether it can ever do the task. Neither answers "will it do this every time." On their 507 workflows, each run 20 times from a clean backend, Claude Opus 5.5 led the overall pass@1 they report, at 67.16%. It passed the same number of tasks on all 20 attempts as Claude Opus 5: 241. A small bump in the headline score bought no extra dependability. Kimi-K3 solved 476 of 507 tasks at least once and only 68 of them every time.

If you pick a model for work that touches real records off the "solved at least once" column, you are buying a demo.

Where the failures sit matters for the week you spend. They label roughly four in five as tool handling, not reasoning — tool usage at 79.9% in an unweighted average of per-model shares, not a unique cause and not a law. Agents usually get far enough to try the workflow, then fail to recover from tool errors, failed preconditions, or empty lookups. That is a retry policy and a smaller tool surface before it is a shopping trip for a newer model.

Cheap Twin still applies: do not send every ticket to the flagship, and do not route on a single-pass score. In their list-rate estimate — not an invoice — the cheapest single success was not the cheapest task that passed every repeat.

Five lines before the next closer ships

  • Required end state is a field diff, written down before the run
  • Post-write readback from the system of record, not the assistant message
  • Customer-visible send stays closed on mismatch
  • Receipt names the fields that matched, the fields that did not, and any extra writes
  • Ship Gate uses a repeat count you can defend — and you say whether you mean best-of-k or every-of-k

Twenty is their number, not a religion. A refund that moves money may need a stricter every-time bar than a draft note. Name k. Say which one you mean. Grade the sentence only when the requirement has no field.

How this sits with the stack

Receipt Gate already refuses the next write when the last step left mush. This tightens the receipt: it has to agree with a readback. Golden Trace should freeze the end state, not only the happy transcript — otherwise a prompt bump that still "sounds resolved" sails through. Smoke evals need one trap where every tool returns ok and a required field is wrong. Permission Envelope still wraps the irreversible send; a correct record does not grant a standing right to email the customer. Cheap Twin can pick the brain. It cannot pick the definition of done.

The test

Take one workflow that writes a record and then tells a human it is finished. Run it five times from the same starting state. After each run, ignore the last message. Read the fields you claimed were the goal.

If any run says resolved while a required field is wrong, your eval is grading the claim.

Sources

Numbers and the retail walkthrough are from the October 3, 2026 joint post. I did not re-run the benchmark. Cost remarks there are the authors' list-rate estimates (OpenRouter snapshot noted as September 20, 2026), not production invoices. They say they have not measured the lift from terminal-state checks, a smaller tool surface, or human approval on this benchmark.

Closing

I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.

Follow for practical AI architecture that ships.

Connect on LinkedIn: Wasim Sheikh · Site: sheikhwasim.com · Notes: Practical AI Notes · X: @anciwasim

Related: Receipt Gate · Golden Trace · Smoke Evals · Ship Gate · Cheap Twin · Permission Envelope · The 5-Layer Agent Stack