01 Compare the pilot setup with the original demo
Sebastian Mueller, a founding partner of MING Labs, described two client pilots that had reached an awkward point: they worked, yet nobody was happy with the quality.[S4][S1] Documents uploaded. Questions received answers. The output simply looked worse than what the team had seen in the earlier demonstration.
His explanation was concrete. The demo had used a frontier model on the vendor’s cloud with external internet access. The pilot used an IT-approved setup with different hosting, a smaller model and no external internet. The prompts and design were similar; the conditions were not.[S1]
That account matters before a buying decision. A sponsor can remember the result on screen while the delivery team is planning a different system. Both may think they are discussing the same agent. The mismatch only becomes visible when someone tries to use its work.
These are Sebastian’s reported observations, not a controlled comparison. They do not tell us how much quality changed or which difference caused it. They do give a project team a useful question: what, exactly, did the demonstration prove about the system we will run?
02 Test with the model and permissions planned for production
Before the next demonstration, ask the delivery team to name the model and version, where it runs, which information it can read and which actions it is allowed to take. Put the intended production setup beside it. If something differs, make that difference visible before discussing results.
A research agent with web access can attempt a different task from one restricted to an internal document set. That is an illustrative distinction, not a verdict that one setup is better. The right question is whether the approved setup can do the agreed job. Broader access is not an acceptance criterion by itself.
In a separate post, Sebastian proposes looking at individual data flows when deciding where model requests may go.[S2] For an acceptance test, turn that into a practical preparation step: have the responsible IT, security and privacy owners establish which inputs, tools and destinations are allowed. Then test inside those boundaries. A convincing demo is no reason to relax them.
If the intended environment is unavailable, show the capability example with its dependencies clearly labelled. Leave readiness for that environment open. A demonstration can inform a pilot without being sufficient evidence to approve production.
03 Agree what would count as a usable result
Start with the person who will receive the work. For an illustrative research role, “produce a briefing” is too loose. The recipient might need source links, unresolved contradictions and a draft delivered into an existing review queue. Write down those requirements before judging the prose.
Use this MING checklist to prepare the acceptance decision. The proposed checks separate what was shown from what still needs proof; they are not a certified standard or a reported result from the client cases.
| What to agree | Evidence to ask for | If it is missing |
|---|---|---|
| Model and hosting | The actual model/version and hosting used in both demo and intended deployment | Record the difference; test the intended configuration before claiming readiness |
| Information and access | Approved inputs, data destinations, tools and permissions for the real task | Resolve the access or scope question with its owner |
| Usable work | Representative tasks with written criteria and an identified person who accepts the result | Clarify what the recipient needs before scoring output |
| Delivery and exceptions | A result reaching its intended review destination; a checked handover for missing evidence or denied access | Keep that part of the work under human control until the handover is verified |
| Effort and variability | Repeated task results, failures, waiting time and human review effort, with the number of attempts | Keep the uncertainty visible; do not substitute a selected best run |
| Acceptance and change | A named decision owner, agreed release conditions and a plan to rerun checks after changes | Leave the unproven scope out of the production commitment |
The checklist becomes useful when each row has an owner and a piece of evidence. “We will handle permissions later” is an open dependency. “The agent can read these approved folders, and the blocked-folder case reached the reviewer” is something the team can inspect.
04 Watch the work finish, including the awkward cases
Anthropic’s evaluation guidance treats the model and surrounding agent system as a combined unit. It also separates an agent saying it completed a task from the actual final outcome, and recommends repeated trials because runs vary.[S3]
Apply that distinction to the research example. Check whether the briefing reached the correct review queue, whether its sources support its claims and whether the reviewer can use it. Then try a missing document, conflicting evidence and a denied permission. These are proposed test cases, not incidents from Sebastian’s clients.
Record every attempt in the chosen set, including unusable results and human corrections. Decide with the task owner how much coverage the consequences of failure require. A tidy percentage without its task set and denominator cannot tell a buyer what was tested. Our operating-cost guide explains why review and correction time belong in the decision too.
05 Approve a bounded role, then keep checking it
The outcome need not be a choice between launching everything and abandoning the project. You might approve a narrower role, retain human approval for a difficult step or resolve a missing integration before testing again. Make that decision against the criteria you agreed, rather than against the memory of the best demo.
Keep the accepted configuration and task set with the handover. When a model, permission or data source changes, rerun the relevant checks. Define who can pause the role if the work stops meeting its standard; our answer on accountability for agent mistakes covers that responsibility.
The useful demo shows what the agent can do under the conditions your team will actually give it.
Bring MING the task, the demo’s setup and the constraints of the intended environment. We can help turn them into a role and acceptance plan for your Hybrid Organisation . Discuss your pilot and its acceptance criteria .
What this guide establishes
- Reported experience: Sebastian describes two client pilots and a demo-to-pilot quality gap. Public task counts, scores, raw logs and results after a fix are unavailable. The account is not a failure-rate study or proof that a particular model or host caused the gap.[S1]
- A proposed method: the checklist and research examples are MING’s editorial synthesis, reviewed on 3 October 2026. They do not promise a measured improvement or cover every production risk.
- Data boundaries: the routing post informs a preparation question. Its broader compliance and performance claims are not adopted. Required approvals remain with the responsible specialists.[S2]
- Source dates: LinkedIn displayed rounded ages. The source list records retrieval dates rather than an inferred publication day. The original meeting illustration is AI-generated and is not used as evidence of a client meeting.