01 Seven correct drafts hid a reduction in errors
You have paid for another round of fixes and want to know whether the agent is doing a better job. The report still shows seven fully correct results. Before deciding that the work has stalled, look at what happened to the other cases.
We met that problem while building an agent that prepares order drafts for an industrial distributor in Germany. It reads incoming orders and matches them to the catalogue; people remain responsible for booking them. Across three rounds of changes, the same 20 test orders still produced seven fully correct drafts. Yet wrong drafts fell from seven to four, and correct escalations rose from six to nine. Our public account places the work in June 2026.[S1][S2]
The share of fully correct drafts was therefore 35% before and after. That calculation is accurate, but it leaves out a change the order desk needs to understand: fewer cases ended with a wrong answer. A manager watching only the 35% could miss that progress. A manager watching only the rising escalations could miss the extra work reaching the team.
Both need the same view of the work. Keep the three counts together, then ask what each result means for the person who receives it.
02 A correct handover must reach someone who can deal with it
In the order example, “fully correct” means every required field in the draft is right without a human correcting it. “Correctly escalated” means the agent recognises what it cannot settle and passes the case, with its source, to the right person. “Wrong” means it makes a decision that does not meet the agreed criteria.[S1]
Consider an illustrative order with an unclear customer number. Guessing a plausible number risks creating the wrong record. Passing the original order and the unresolved question to the responsible colleague preserves a chance to resolve it properly. That is a useful distinction even when neither route produces an immediately usable draft.
But a message saying “escalated” is not enough. Check who received it, whether the original order is attached and whether it arrived in time for the work to continue. If it sits in an unattended queue, the handover has not done its job. Record that failure rather than counting it as a successful escalation.
Nor should an agent pass every ordinary case back to people. Read the reasons for the growing middle group. Some may reveal appropriate caution; others may reveal missing data, unnecessary restrictions or a task the agent still cannot handle. The count tells you where to look. It does not make the decision for you.
03 A higher score was not worth two additional wrong-data drafts
Our tooling suggested lowering a confidence threshold to recover more clean drafts. We simulated that change first. The reported result was four additional clean drafts and two additional drafts carrying wrong data. The proposal was rejected because the project had already set a firm boundary against silently introducing wrong data.[S1]
Those were simulated outcomes. We did not ship that threshold change, and the figures do not describe two incorrect orders booked in production. They explain why a better-looking headline would have been a poor release decision.
Before changing a setting, write down the trade you are willing to accept. Which errors must stop the release? Who checks cases that remain uncertain? Who has the authority to expand what the agent may do? Keep that agreement beside the results so that a numerical improvement cannot quietly replace it.
04 Compare the same tasks and inspect what never reached the agent
For the next test, keep the inputs, acceptance criteria and operating conditions explicit. A different mix of orders can move the score without telling you whether the change helped. Anthropic’s evaluation guidance also recommends repeated trials and checking the actual final outcome, rather than relying on the agent’s account of what it did.[S3]
The three groups cover the assessed cases, not everything the business expected to happen. Reconcile expected incoming work with what reached the agent. Keep missing inputs, failed runs and cases awaiting assessment visible in their own records. Do not drop them from a denominator or quietly call them correct. Our audit-trail guide explains how to connect a task to evidence of what actually happened.
Then look at the human side of the result. How long do escalated cases wait? How many minutes does a person spend checking and correcting the drafts? Does the finished work arrive when it is needed? Record those measures alongside the counts. Our cost guide shows why accepted results and human effort belong in the same operating decision.
| Keep visible | Define before comparing |
|---|---|
| Fully correct, correctly escalated, wrong | Counts and shares of the same assessed test set; criteria for each category. |
| Missing or unfinished work | Expected incoming tasks, tasks received and cases still awaiting assessment. |
| Human review effort | Minutes per case, including corrections; identify which cases the measure covers. |
| Waiting and completion | Time from receipt to the agreed result, plus cases still waiting at the cut-off. |
05 Review one set of cases with the person who accepts the work
Choose a documented set of recent cases and ask the process owner to classify them with you. Define what counts as a correct result before scoring. Inspect every alleged handover in that set and keep unresolved judgements open. The source suggests starting with a hundred cases; that is an exercise, not a universal sample size or a promise that an afternoon will be enough.[S1]
Use what you find to choose the next change: repair a data gap, improve a handover, tighten a boundary or test a capability again. Our acceptance guide helps turn that choice into a bounded test. If you want MING to help, bring the task, the three counts and an example the team disagrees about to a Hybrid Organisation discussion .
What the comparison establishes
- The 20-order comparison is published MING experience from one engagement. Its underlying logs are private; this page does not independently reproduce them. The post and article belong to the same source family.
- Counts concern drafts. The separate 115-order sample in the source is not the denominator here. Aggregate counts do not identify which individual cases changed category.
- The proposed scorecard and review steps do not prove a general reliability level, time saving or readiness for more autonomy. More escalations are useful only when they serve the task and the team can handle them.