Answer

How do you test an AI agent before it goes into production?

Test the whole agent in the environment you intend to run: the approved model, hosting, data and permissions. Agree what a usable result looks like, then check representative tasks, exceptions and human handover against that standard. Sebastian Mueller’s reported demo-to-pilot gap shows why matching the prompt alone is not enough; it does not establish a universal cause of poor performance.[S1]

Sebastian’s reported case

The pilot used a smaller model without external internet access

The demo

  • Frontier model
  • Vendor cloud
  • External internet available

The pilot

  • Smaller model
  • Approved hosting
  • No external internet

Test the conditions the work will actually depend on.

Reported differences between demo and pilot, not a measured performance comparison. Include model, hosting and access in acceptance. Use the acceptance checklist.
Last updated: | Next review: Acceptance guide Machine-readable record

By MING Labs · Editorial team

01 Compare the pilot setup with the original demo

Sebastian Mueller, a founding partner of MING Labs, described two client pilots that had reached an awkward point: they worked, yet nobody was happy with the quality.[S4][S1] Documents uploaded. Questions received answers. The output simply looked worse than what the team had seen in the earlier demonstration.

His explanation was concrete. The demo had used a frontier model on the vendor’s cloud with external internet access. The pilot used an IT-approved setup with different hosting, a smaller model and no external internet. The prompts and design were similar; the conditions were not.[S1]

That account matters before a buying decision. A sponsor can remember the result on screen while the delivery team is planning a different system. Both may think they are discussing the same agent. The mismatch only becomes visible when someone tries to use its work.

These are Sebastian’s reported observations, not a controlled comparison. They do not tell us how much quality changed or which difference caused it. They do give a project team a useful question: what, exactly, did the demonstration prove about the system we will run?

02 Test with the model and permissions planned for production

Before the next demonstration, ask the delivery team to name the model and version, where it runs, which information it can read and which actions it is allowed to take. Put the intended production setup beside it. If something differs, make that difference visible before discussing results.

A research agent with web access can attempt a different task from one restricted to an internal document set. That is an illustrative distinction, not a verdict that one setup is better. The right question is whether the approved setup can do the agreed job. Broader access is not an acceptance criterion by itself.

In a separate post, Sebastian proposes looking at individual data flows when deciding where model requests may go.[S2] For an acceptance test, turn that into a practical preparation step: have the responsible IT, security and privacy owners establish which inputs, tools and destinations are allowed. Then test inside those boundaries. A convincing demo is no reason to relax them.

If the intended environment is unavailable, show the capability example with its dependencies clearly labelled. Leave readiness for that environment open. A demonstration can inform a pilot without being sufficient evidence to approve production.

03 Agree what would count as a usable result

Start with the person who will receive the work. For an illustrative research role, “produce a briefing” is too loose. The recipient might need source links, unresolved contradictions and a draft delivered into an existing review queue. Write down those requirements before judging the prose.

Use this MING checklist to prepare the acceptance decision. The proposed checks separate what was shown from what still needs proof; they are not a certified standard or a reported result from the client cases.

What to agreeEvidence to ask forIf it is missing
Model and hostingThe actual model/version and hosting used in both demo and intended deploymentRecord the difference; test the intended configuration before claiming readiness
Information and accessApproved inputs, data destinations, tools and permissions for the real taskResolve the access or scope question with its owner
Usable workRepresentative tasks with written criteria and an identified person who accepts the resultClarify what the recipient needs before scoring output
Delivery and exceptionsA result reaching its intended review destination; a checked handover for missing evidence or denied accessKeep that part of the work under human control until the handover is verified
Effort and variabilityRepeated task results, failures, waiting time and human review effort, with the number of attemptsKeep the uncertainty visible; do not substitute a selected best run
Acceptance and changeA named decision owner, agreed release conditions and a plan to rerun checks after changesLeave the unproven scope out of the production commitment

The checklist becomes useful when each row has an owner and a piece of evidence. “We will handle permissions later” is an open dependency. “The agent can read these approved folders, and the blocked-folder case reached the reviewer” is something the team can inspect.

04 Watch the work finish, including the awkward cases

Anthropic’s evaluation guidance treats the model and surrounding agent system as a combined unit. It also separates an agent saying it completed a task from the actual final outcome, and recommends repeated trials because runs vary.[S3]

Apply that distinction to the research example. Check whether the briefing reached the correct review queue, whether its sources support its claims and whether the reviewer can use it. Then try a missing document, conflicting evidence and a denied permission. These are proposed test cases, not incidents from Sebastian’s clients.

Record every attempt in the chosen set, including unusable results and human corrections. Decide with the task owner how much coverage the consequences of failure require. A tidy percentage without its task set and denominator cannot tell a buyer what was tested. Our operating-cost guide explains why review and correction time belong in the decision too.

05 Approve a bounded role, then keep checking it

The outcome need not be a choice between launching everything and abandoning the project. You might approve a narrower role, retain human approval for a difficult step or resolve a missing integration before testing again. Make that decision against the criteria you agreed, rather than against the memory of the best demo.

Keep the accepted configuration and task set with the handover. When a model, permission or data source changes, rerun the relevant checks. Define who can pause the role if the work stops meeting its standard; our answer on accountability for agent mistakes covers that responsibility.

The useful demo shows what the agent can do under the conditions your team will actually give it.

Bring MING the task, the demo’s setup and the constraints of the intended environment. We can help turn them into a role and acceptance plan for your Hybrid Organisation . Discuss your pilot and its acceptance criteria .

What this guide establishes

  • Reported experience: Sebastian describes two client pilots and a demo-to-pilot quality gap. Public task counts, scores, raw logs and results after a fix are unavailable. The account is not a failure-rate study or proof that a particular model or host caused the gap.[S1]
  • A proposed method: the checklist and research examples are MING’s editorial synthesis, reviewed on 3 October 2026. They do not promise a measured improvement or cover every production risk.
  • Data boundaries: the routing post informs a preparation question. Its broader compliance and performance claims are not adopted. Required approvals remain with the responsible specialists.[S2]
  • Source dates: LinkedIn displayed rounded ages. The source list records retrieval dates rather than an inferred publication day. The original meeting illustration is AI-generated and is not used as evidence of a client meeting.

A proposed MING acceptance method based on reported practitioner experience and technical guidance. No measured recovery, benchmark or compliance assurance.

Sources

[S1]
Two clients this week, same wall. The pilot works. Nobody is happy. Sebastian Mueller on LinkedIn · Accessed 2026-10-03 Supports: Two reported client pilots that answered questions but appeared worse than earlier demos, Reported differences in model, hosting and external internet access; proposal to demonstrate the intended environment Public practitioner account and author comment. LinkedIn displayed 1w. No public task set, scores, raw logs or measured recovery. Two cases are not a failure rate.
[S2]
We built a pilot to the strictest reading of European data protection. Sebastian Mueller on LinkedIn · Accessed 2026-10-03 Supports: Proposal to classify individual data flows before deciding how to route model requests Public proposal and author comment; LinkedIn displayed 1w. The graphic proposes routing by data category. Its compliance and comparative performance claims are not relied on here.
[S3]
Demystifying evals for AI agents Anthropic · 2026-01-09 Supports: Evaluate the model and surrounding agent system together, Distinguish the final outcome from a statement that work was completed; repeat trials because runs vary
[S4]
About MING Labs: founders MING Labs · Accessed 2026-10-03 Supports: Sebastian Mueller is a founding partner of MING Labs Undated company page; date records retrieval.

Frequently asked questions

Does a weaker pilot mean we need a bigger model?
Not necessarily. Check what changed: model, available information, tools, permissions or the task itself. The cases on this page do not isolate those factors. Test the relevant difference under approved conditions before attributing the gap to model size.
What if the production environment is not ready for the demo?
Label the demonstration as a capability example and list the conditions it depends on. Treat performance in the intended environment as unverified until you can test it. Do not use an unrestricted demonstration as acceptance evidence for a restricted deployment.
How many test cases are enough?
There is no universal number in the sources here. Choose coverage with the task owner according to the work and the consequences of failure. Include ordinary tasks, meaningful exceptions and blocked access, record the denominator, and repeat important cases. A small pilot cannot prove every future case will work.
Does European hosting make the agent compliant?
Hosting location alone does not establish compliance. The relevant data flows, providers, access, contracts and purpose need review by the responsible specialists. This page provides a testing method, not a legal assessment or permission to move data.
All Insights