AI case studies often compress a complicated intervention into one sentence:
The company introduced AI and reduced processing time by 60%.
The sentence may be accurate. It may also leave out the information needed for your decision:
- What was measured?
- What was the starting point?
- Did the workflow change at the same time?
- Were difficult cases excluded?
- Did employees perform hidden review or cleanup?
- Was the result achieved in a pilot or normal operations?
- Did quality, cost or customer outcomes change?
- Would the result survive another team, dataset or volume?
A case is useful when the reader keeps the claim inside the boundary of its evidence. It becomes dangerous when a local observation is presented as a causal and transferable promise.
The short answer
Read every AI case at five levels:
- Existence: a system was built or deployed.
- Capability: it produced a defined output on selected examples.
- Process outcome: an operational measure changed.
- Business outcome: cost, risk, revenue or customer experience changed.
- Attribution and transfer: the intervention caused the change and the mechanism is likely to work in another context.
Evidence at one level does not automatically prove the next.
A working demo proves existence. A test-set score provides evidence about a defined capability. A faster pilot shows an observed process change. None of these, alone, proves a durable business result or that another organisation will obtain the same effect.
Outcome and attribution are different questions
Monitoring can show that something changed. Attribution asks whether the AI intervention caused that change.
The UK government’s Quality in Policy Impact Evaluation guidance makes the distinction explicit: observing an outcome is not sufficient to establish that the intervention was responsible. A comparison or a credible explanation of the counterfactual is needed.
Small businesses do not need an academic study for every workflow. They do need language that matches the strength of the evidence.
Compare:
- “Response time fell during the four-week pilot.”
- “The AI system reduced response time.”
- “The AI system will reduce response time in your business.”
The first is an observation. The second is causal. The third adds a transfer claim. Each requires more evidence than the previous statement.
The AI Case Evidence Card
Use this card before citing a case in a proposal or using it to approve a project. It is a Methodfield working tool, not a validated research-quality rating.
1. Source and interest
Record who published the case and who benefits from the claim.
Possible source types include:
- the organisation using the system;
- the vendor selling the system;
- a joint customer story;
- a regulator or public body;
- an independent evaluation;
- a media summary of another source.
A vendor case can still contain useful operational evidence. It should be labelled vendor-reported and read with the missing incentives and methods in mind.
2. Context and boundary
Record:
- organisation and business model;
- country and regulatory context;
- workflow and users;
- data type;
- sample or volume;
- pilot or production status;
- time period;
- included and excluded cases.
Without this boundary, the reader cannot assess relevance.
3. Baseline and comparison
Ask what would have happened without the intervention.
Useful comparisons include:
- the same process before deployment;
- a parallel team or location;
- randomly assigned cases;
- a staged rollout;
- a defined business-as-usual process;
- repeated measures before and after a change.
A before-and-after comparison is often practical, but it remains vulnerable to other changes occurring at the same time.
4. Intervention
Describe what actually changed.
“Introduced AI” is not enough. The case should distinguish:
- model or service;
- prompt, rules and knowledge sources;
- integrations;
- human review;
- process redesign;
- training and support;
- changed staffing or service hours;
- fallback and exception handling.
This reveals the mechanism. It also prevents the model from receiving credit for a broader operational redesign.
5. Measure
Capture the metric definition, not only the number.
For example, “processing time” may mean:
- model execution time;
- employee touch time;
- elapsed time from arrival to completion;
- time excluding escalated cases;
- time to draft rather than time to an approved outcome.
The measurement window, denominator and treatment of failures change the meaning of the result.
6. Quality and unintended effects
Efficiency should be read beside:
- correction and rework;
- false positive and false negative consequences;
- escalations;
- complaints;
- inconsistent outcomes between groups;
- user workarounds;
- new risks or delays downstream.
A faster workflow that moves more errors to the next step has not necessarily improved.
7. Resources and operating cost
Record resources added only for the pilot:
- expert reviewers;
- manual data preparation;
- vendor support;
- temporary integrations;
- free credits;
- unusually narrow scope;
- selected users or cases.
Separate build cost, recurring cost and human oversight. A technically successful case may still have an unfavourable operating model.
8. Replication and transfer boundary
Ask:
- Was the result repeated?
- Did it hold as volume increased?
- Did new users achieve it?
- Which data, skills and process conditions were essential?
- Which part is likely to transfer: the technology, the workflow pattern or only the question being tested?
Often the most transferable element is not the reported percentage. It is the architecture: extract, validate, route, review and log.
A synthetic example
Suppose a vendor reports that a retailer used AI to review product information and cut preparation time by 60%.
The case may justify the following statement:
In the reported workflow and measurement period, the combined AI and review process completed the selected product-information task faster than the previous approach.
It does not automatically justify:
- AI caused the entire improvement;
- product information became more accurate;
- labour cost fell by 60%;
- the result applies to every product category;
- the workflow will produce the same result in another retailer;
- the system can release information without human review.
The correct next action is not to reject the case. It is to borrow the pattern and design a local test around the missing questions.
Four decisions after reading a case
Borrow the pattern
Use when the workflow mechanism is relevant but the reported result is not transferable. Convert the case into a local hypothesis.
Replicate locally
Use when the expected value is meaningful and the largest uncertainty can be tested with representative cases, a baseline and a controlled boundary.
Scale cautiously
Use when the result has survived repeated use, normal operating conditions and defined quality gates. Expansion should still preserve monitoring and a stop condition.
Reject the inference
Use when the public claim goes beyond its evidence, the context is materially different or the intervention cannot be reconstructed well enough to test.
Rejecting the inference does not mean the technology is useless. It means the case does not support the proposed decision.
Design the prototype to create evidence
A useful prototype should answer four classes of questions.
Capability
Can the system perform the bounded task on representative cases?
Workflow
Does it reduce elapsed time, touch time, missing information or rework once review and exception handling are included?
Consequence
What happens when the system is wrong, uncertain or unavailable?
Business effect
Does the changed workflow improve the outcome that justified the project, at an acceptable operating cost?
The UK guidance on evaluating AI interventions recommends establishing a baseline early, documenting business as usual, considering a comparison and adapting evaluation as the system evolves. Those principles are useful even when a small business applies them proportionately rather than running a formal impact study.
How this connects to existing Methodfield content
The guide From Observation to Action defines evidence gates before authority grows. This article clarifies what the evidence at those gates should and should not claim.
AI Automation Should Improve Quality shows how AI can inspect a process rather than only produce output. The same separation is important in evaluation: a system should not generate an output, declare it correct and count the result as independent quality evidence.
Common evidence mistakes
- quoting a percentage without its denominator or time window;
- treating vendor-reported evidence as independent evaluation;
- calling a selected pilot “production”;
- measuring model speed instead of complete workflow time;
- excluding review and corrections from operating cost;
- presenting correlation as causation;
- assuming a result transfers because the industry label is the same;
- hiding failed or escalated cases;
- changing the metric after seeing the result;
- using testimonials as a substitute for objective evidence.
Final position
An AI case is not a promise. It is a bounded piece of evidence about a system, a workflow and a context.
Use public cases to discover patterns and questions. Use a local prototype to test the mechanism. Use production monitoring to learn whether the result survives real work.
The honest conclusion is often narrower than the headline—and much more useful for a decision.
Sources
- HM Treasury and Evaluation Task Force, Guidance on the Impact Evaluation of AI Interventions (opens in a new tab), updated 15 May 2026.
- HM Treasury and Evaluation Task Force, The Magenta Book (opens in a new tab), updated 15 May 2026.
- UK Government, Quality in Policy Impact Evaluation (opens in a new tab), updated 15 May 2026.
- NIST, AI Risk Management Framework Core (opens in a new tab), version available on 26 August 2026.
- US Federal Trade Commission, Advertising FAQs: A Guide for Small Business (opens in a new tab), version available on 26 August 2026.
