You prove it by testing it, on a schedule, with a consequence attached. Not with a training certificate, not with documentation, and not with a human-in-the-loop approval step that a tired person clicks through. In MING Labs’ own fleet the test is written into the agent’s job description: the operator must reproduce that agent’s three primary outputs unaided, within 30 minutes, once a quarter.[S3]
01The thing you are actually trying to catch
The failure is quiet. Work keeps shipping, quality looks fine, and the ability to do that work drains out of the team without anyone noticing, because nothing breaks on the day it leaves. It only surfaces later, when the output is challenged, the model changes, or the agent is decommissioned, and nobody in the room can rebuild what it produced.
This is measurable, and someone outside our lane has measured it. In a randomised controlled trial with 52 engineers learning an unfamiliar library, the group using AI assistance finished in roughly the same time as the control group and scored 17 points lower on a comprehension test afterwards, 50 percent against 67 percent.[S1] The work arrived. The understanding did not. Addy Osmani calls the result comprehension debt: the codebase looks healthy while comprehension hollows out underneath it.
His frame is scoped to the system. Ours is scoped to the person, because an organisation does not lose the ability to defend its work in the abstract. A named human loses it.
02The test: the Comprehension Obligation
The Comprehension Obligation is the clause we put in an agent’s job description. The human operator must stay able to reproduce that agent’s primary outputs unaided, within 30 minutes, rehearsed each quarter. Fail the rehearsal, and the agent drops one autonomy level, from acting-with-log back to drafting-for-review, and stays there until the operator is current again.[S3]
Two properties do the work here. It is a capability test, not a record: you cannot pass it by having attended something. And failing it costs the agent something, which is what stops the rehearsal from becoming a formality. An obligation with no consequence is a sentiment.
The takeaway: most answers to this question prove exposure. Only one proves capability.
| What a company usually does | What it actually proves | What it misses |
|---|---|---|
| AI training course, certificate on file | The person was exposed to the material | Whether they can still do the work the agent now does |
| Documentation of what the agent does | The process is written down | Whether any human could execute it unaided |
| Human-in-the-loop approval step | Someone clicked approve | Whether the approver understood what they approved |
| Comprehension Obligation | A named human reproduced the agent’s output, unaided, in 30 minutes, this quarter | It tests reproduction, not judgment quality. It is a floor, not a ceiling. |
03Why the law creates this question and does not answer it
Article 4 of the EU AI Act has required providers and deployers to ensure a sufficient level of AI literacy among their staff since 2 February 2025. What it does not do is say how. The obligation is assessed proportionately and case by case, and no method, threshold, or test is prescribed anywhere in it. Supervision was also, until now, nobody’s job: enforcement sits with national market surveillance authorities, and they begin supervising on 2 August 2026.[S2]
So a duty that was easy to defer for eighteen months is about to be examined by someone, and the text that creates it offers no way to demonstrate you have met it. That gap is not a reason to buy a compliance product. It is a reason to have a test you can show, and to have been running it before anyone asked.
To be exact about our own lane: the Comprehension Obligation is not a compliance instrument and we are not lawyers. It tests one thing, whether the named operator of an agent can still reproduce that agent’s work. That is narrower than what Article 4 asks for and considerably harder than what most organisations will offer in its place.
04The measurement, scoped
The rule, as a countable: three primary outputs, 30 minutes, unaided, once per quarter, one autonomy level at stake.
Measurement scope
- Sample: MING Labs’ own production fleet, seven agents under named human owners, as of July 2026.
- Rollout status: the clause is live in the Chief of Staff agent’s job description and is rolling across the remaining agents. It is not yet in all seven.[S4]
- Measured: whether the named operator can reproduce that agent’s three primary outputs unaided, inside 30 minutes, without the agent.
- Consequence: a failed rehearsal demotes the agent one autonomy level until the operator is current again.
- Not measured, and therefore not claimed: we do not publish a rehearsal pass or failure rate. The rollout is incomplete, so any rate we quoted would be a number without a denominator.
- Disclosure: scope only. Internal rehearsal protocols, prompts, and per-agent tooling are not published.
The obligation is the human half of a rule we already apply to the machine half: an agent may not grade its own work, because verification is a job, not a setting . The Comprehension Obligation is the same principle pointed at the operator. It is what keeps judgment with people inside a hybrid organisation while the agents do the work, and it enforces the layer of the ABC Framework that never leaves a human: judgment and accountability.
The full definition, its origin, and how it sits inside an agent’s job description are on the concept page for the Comprehension Obligation .