Answer

When should you shut down an AI agent?

Week-one polish is noise. The Acted-On Rate, the share of agent output a human acts on, is the shutdown signal.

Shut down an AI agent when its Acted-On Rate stays at zero: nobody uses or acts on its output, however busy it looks. Volume is not the signal; use is. MING Labs shut down its first agent, Major Tom, after 13 days and 764 messages because not one output changed a decision. A slow colleague improves that rate week over week. Noise does not.

Last updated: July 2026 | Next review: January 2027 Proprietary evidence Machine-readable record
Based on articleWe fired an AI agent after 13 days

Shut down an AI agent when nobody acts on its output. Not when it makes mistakes, and not when it is slow to find its feet, but when its work changes nothing: no draft gets sent, no analysis moves a decision, no task gets picked up by a colleague. MING Labs runs its production agent fleet on this criterion and used it to shut down its first agent after thirteen days.[S1]

01The signal: the Acted-On Rate

The Acted-On Rate is the share of an agent’s output that a named human actually uses or acts on. A sent draft counts. A forwarded analysis counts. A report that shapes a decision counts. Output that is technically correct and practically ignored does not count, no matter how much of it there is.[S2]

Activity metrics flatter agents. Message volume, tasks completed, and automation rate all measure motion, and an agent can score high on every one of them while producing nothing anyone touches. The Acted-On Rate is the one number in MING’s fleet reviews that separates a slow-starting colleague from noise.

The takeaway: judge the trend of use, not the volume of output. The two failure profiles look identical on activity metrics and opposite on the Acted-On Rate.

SignalSlow-starting colleagueNoise: shut it down
Output volumeLow to normal, unevenHigh and steady
Acted-On RateAbove zero by week two, climbingFlat at zero
Human responsePeople correct and resend its draftsPeople ignore its messages
Role designOwned domain, named human ownerCapability without a domain
What to doKeep it, tighten the roleShut it down, redesign the role

02What Major Tom taught us

MING Labs gave its first agent, Major Tom, a capability: coordination across the founder team. It never gave him a domain. Over thirteen days he sent 764 messages, and the team acted on none of them. Every message was technically correct and practically useless. We shut him down on March 30 and published the audit.[S1]

The lesson was not that the model was weak. The lesson was that the role was wrong: capability without accountability is noise. The agents that replaced him got job descriptions, owned domains, and a named human owner, and their output started being used within days. Shutting an agent down is not the opposite of adopting agents. It is how a fleet stays honest.

Nor was he unique. Across MING’s fleet, 15 to 20 percent of agents are decommissioned within their first quarter.[S2] That rate, why it is a sign of health rather than failure, and the retirements behind it have their own page: how often AI agents get shut down .

03Why week one tells you almost nothing

The obvious objection is that any new colleague needs time, and killing an agent in week one punishes slow starters. That is true for quality and irrelevant for use. Early output is allowed to be mediocre; a human editing a mediocre draft is still acting on it, and that counts. What week one cannot excuse is a rate that never leaves zero while volume climbs, because that pattern does not improve with polish. It improves with a different role.

The measure also has to sit outside the agent. An agent asked to grade its own usefulness will find itself useful, which is why verification is a job, not a setting . In MING’s fleet the Acted-On Rate is read from what humans did, never from what the agent reports about itself.[S2]

04The measurement, scoped

The founding data point: 13 days, 764 messages, 0 acted-on outputs.

Measurement scope

  • Period: 17 to 30 March 2026, ended by shutdown.
  • Sample: one internal agent (Major Tom, fleet coordination role) in MING Labs’ own production fleet; message audit n=764.
  • Measured: whether any output was used or acted on by a named human.
  • Success: output used at least once, rate climbing. Shutdown trigger: Acted-On Rate flat at zero across the window.
  • Disclosure: scope only. Internal prompts, protocols, and per-agent tooling are not published.

Since then the same measure has governed promotions, not only shutdowns. Agents in the fleet earn autonomy levels as their Acted-On Rate climbs, each under a named human owner who must pass the Comprehension Obligation to keep the autonomy level up.[S2] Which work an agent may own in the first place is a separate question, and the ABC Framework answers it: judgment stays human, structured expert work is shared, routine gets owned.

Sources

[S1]
We fired an AI agent after 13 days (Major Tom message audit, n=764)MING Labs · 2026-03-30 Supports: Major Tom shut down after 13 days and 764 messages, zero acted-on output as the shutdown reason, capability without accountability diagnosis
[S2]
MING Labs operating record: fleet task reviews and autonomy levelsMING Labs (internal) · 2026-07-01 Supports: Acted-On Rate as the fleet's standing shutdown and promotion signal, healthy agents show acted-on output within the first two weeks, autonomy levels promoted and demoted under named human ownership, 15 to 20 percent of agents decommissioned within their first quarter

Frequently asked questions

Isn't two weeks too early to judge an AI agent?
Two weeks is too early to judge quality. It is not too early to judge use. A slow-starting colleague produces imperfect output that people still correct, forward, or build on. If nothing has been acted on at all by week two, the role is wrong, and more polish will not fix it.
What Acted-On Rate should a healthy agent show?
There is no universal benchmark, and MING Labs does not publish one. The signal is the trend: acted-on output should appear within the first two weeks and climb as trust grows. MING promotes agents through autonomy levels when the rate climbs, and demotes or shuts down when it stays flat.
Is a flat Acted-On Rate the agent's fault or the role's fault?
Usually the role's. Major Tom had a capability, coordination, but no domain and no accountability, so every message was technically correct and practically useless. Shut down the role, redesign it around owned work with a named human owner, then rehire.
How is the Acted-On Rate different from task volume or automation rate?
Volume counts what the agent does. Automation rate counts what it does without help. The Acted-On Rate counts what a human actually uses. An agent can score high on the first two and still be noise, which is exactly what 764 unused messages look like.
All Insights