Answer

Which metrics show whether an AI agent is improving?

Count fully correct results, correct handovers to people and wrong results separately on the same test cases. Then check missed inputs, waiting time and human review effort. In our reported 20-order test, wrong drafts fell from seven to four while fully correct drafts stayed at seven. That is useful progress, but it does not establish readiness for autonomous booking.[S1][S2]

Reported MING case · June 2026 · 20 test orders

The same 20 orders produced fewer wrong drafts.

The same 20 test orders before and after:

OutcomeBeforeAfter
Fully correct77
Correctly escalated69
Wrong74
The fully correct share stayed at 7 of 20 (35%). Wrong drafts fell from 7 to 4; correctly escalated cases rose from 6 to 9. These are reported draft outcomes from one engagement, not automatically booked orders. Published project account.
Last updated: | Next review: Published MING experience Machine-readable record

By MING Labs · Editorial team

01 Seven correct drafts hid a reduction in errors

You have paid for another round of fixes and want to know whether the agent is doing a better job. The report still shows seven fully correct results. Before deciding that the work has stalled, look at what happened to the other cases.

We met that problem while building an agent that prepares order drafts for an industrial distributor in Germany. It reads incoming orders and matches them to the catalogue; people remain responsible for booking them. Across three rounds of changes, the same 20 test orders still produced seven fully correct drafts. Yet wrong drafts fell from seven to four, and correct escalations rose from six to nine. Our public account places the work in June 2026.[S1][S2]

The share of fully correct drafts was therefore 35% before and after. That calculation is accurate, but it leaves out a change the order desk needs to understand: fewer cases ended with a wrong answer. A manager watching only the 35% could miss that progress. A manager watching only the rising escalations could miss the extra work reaching the team.

Both need the same view of the work. Keep the three counts together, then ask what each result means for the person who receives it.

02 A correct handover must reach someone who can deal with it

In the order example, “fully correct” means every required field in the draft is right without a human correcting it. “Correctly escalated” means the agent recognises what it cannot settle and passes the case, with its source, to the right person. “Wrong” means it makes a decision that does not meet the agreed criteria.[S1]

Consider an illustrative order with an unclear customer number. Guessing a plausible number risks creating the wrong record. Passing the original order and the unresolved question to the responsible colleague preserves a chance to resolve it properly. That is a useful distinction even when neither route produces an immediately usable draft.

But a message saying “escalated” is not enough. Check who received it, whether the original order is attached and whether it arrived in time for the work to continue. If it sits in an unattended queue, the handover has not done its job. Record that failure rather than counting it as a successful escalation.

Nor should an agent pass every ordinary case back to people. Read the reasons for the growing middle group. Some may reveal appropriate caution; others may reveal missing data, unnecessary restrictions or a task the agent still cannot handle. The count tells you where to look. It does not make the decision for you.

03 A higher score was not worth two additional wrong-data drafts

Our tooling suggested lowering a confidence threshold to recover more clean drafts. We simulated that change first. The reported result was four additional clean drafts and two additional drafts carrying wrong data. The proposal was rejected because the project had already set a firm boundary against silently introducing wrong data.[S1]

Those were simulated outcomes. We did not ship that threshold change, and the figures do not describe two incorrect orders booked in production. They explain why a better-looking headline would have been a poor release decision.

Before changing a setting, write down the trade you are willing to accept. Which errors must stop the release? Who checks cases that remain uncertain? Who has the authority to expand what the agent may do? Keep that agreement beside the results so that a numerical improvement cannot quietly replace it.

04 Compare the same tasks and inspect what never reached the agent

For the next test, keep the inputs, acceptance criteria and operating conditions explicit. A different mix of orders can move the score without telling you whether the change helped. Anthropic’s evaluation guidance also recommends repeated trials and checking the actual final outcome, rather than relying on the agent’s account of what it did.[S3]

The three groups cover the assessed cases, not everything the business expected to happen. Reconcile expected incoming work with what reached the agent. Keep missing inputs, failed runs and cases awaiting assessment visible in their own records. Do not drop them from a denominator or quietly call them correct. Our audit-trail guide explains how to connect a task to evidence of what actually happened.

Then look at the human side of the result. How long do escalated cases wait? How many minutes does a person spend checking and correcting the drafts? Does the finished work arrive when it is needed? Record those measures alongside the counts. Our cost guide shows why accepted results and human effort belong in the same operating decision.

Keep visibleDefine before comparing
Fully correct, correctly escalated, wrongCounts and shares of the same assessed test set; criteria for each category.
Missing or unfinished workExpected incoming tasks, tasks received and cases still awaiting assessment.
Human review effortMinutes per case, including corrections; identify which cases the measure covers.
Waiting and completionTime from receipt to the agreed result, plus cases still waiting at the cut-off.

05 Review one set of cases with the person who accepts the work

Choose a documented set of recent cases and ask the process owner to classify them with you. Define what counts as a correct result before scoring. Inspect every alleged handover in that set and keep unresolved judgements open. The source suggests starting with a hundred cases; that is an exercise, not a universal sample size or a promise that an afternoon will be enough.[S1]

Use what you find to choose the next change: repair a data gap, improve a handover, tighten a boundary or test a capability again. Our acceptance guide helps turn that choice into a bounded test. If you want MING to help, bring the task, the three counts and an example the team disagrees about to a Hybrid Organisation discussion .

What the comparison establishes

  • The 20-order comparison is published MING experience from one engagement. Its underlying logs are private; this page does not independently reproduce them. The post and article belong to the same source family.
  • Counts concern drafts. The separate 115-order sample in the source is not the denominator here. Aggregate counts do not identify which individual cases changed category.
  • The proposed scorecard and review steps do not prove a general reliability level, time saving or readiness for more autonomy. More escalations are useful only when they serve the task and the team can handle them.

The same 20 test orders in one June 2026 engagement. Reported draft outcomes, not a general reliability benchmark or autonomous booking result.

Sources

[S1]
Every fix worked. The accuracy number never moved. Sebastian Mueller, The Hybrid Company / Medium · 2026-09-30 Supports: Reported June 2026 order-draft counts, category definitions, draft-only scope and rejected threshold simulation Published MING experience. Full original and six graphics re-read; underlying engineering records are private, not independently verified.
[S2]
The same 20 test orders, before and after Sebastian Mueller, LinkedIn · Accessed 2026-10-08; displayed 5d Supports: Explicit confirmation that the before/after counts concern the same 20 test orders Full public original and graphic read. Same evidence family as S1, not independent corroboration.
[S3]
Demystifying evals for AI agents Anthropic · 2026-01-09; checked 2026-10-08 Supports: Defined task criteria, repeated trials and checking the actual outcome in the environment Primary engineering guidance, not evidence for the MING client result.

Frequently asked questions

Is a higher escalation rate good or bad?
It depends on what replaced what and whether the handover works. Catching an unsafe guess can be progress. Passing ordinary work back unnecessarily can overload the team. Compare the same tasks, inspect the reason for each escalation and measure whether a person receives and resolves it in time.
Are 20 test orders enough to prove an agent is reliable?
No. The reported set explains one development decision. It does not establish reliability across a production workload. Choose coverage for the task and consequences of failure, include representative and difficult cases, and repeat trials. Keep the test set separate from production monitoring.
Does fully correct mean the order was booked automatically?
Not in this case. The agent prepared order drafts and never booked orders on its own. Fully correct describes the draft’s fields. A production measure must name the actual result being accepted; a draft, a human-approved booking and a delivered order are different outcomes.
Which number should we improve first?
Agree that with the person responsible for the process. Start with unacceptable errors and their consequences, then assess useful output, human effort and time to completion together. On the reported engagement, zero silent wrong data was a pre-agreed gate. It was not a claim that all errors had disappeared.
All Insights