Skip to content
Analysis8 min readSources reviewed

What does AI efficiency actually mean?

Define AI efficiency with accepted quality, elapsed time, full cost, risk and the burden placed on the team.

For Business owners, operations leaders and AI system designers

What does AI efficiency actually mean?

An engineer may mean tokens per second by “efficient”. A finance lead may mean cost per request. A service manager may mean customers helped without a second contact. A worker may ask how much review the tool creates. All are valid observations. None is sufficient on its own.

For an operational workflow, define efficiency as the share of accepted outcomes delivered on time with tolerable risk per unit of total resource. The definition becomes useful only when “accepted” is written down: required fields, valid evidence, permitted actions, editorial quality or another task-specific rule.

Measure the dimensions together

Quality includes completeness, correctness, appropriate refusal and uncertainty. Speed includes time to an accepted result, not only time to the first token. Cost includes models, tools, data, infrastructure, repeats, human review and correction. Risk includes serious mistakes, access violations and actions that cannot be reconstructed. Team burden includes interruptions, exceptions and ongoing maintenance.

Stanford's HELM (opens in a new tab) developed a multi-metric approach because accuracy alone misses properties such as calibration and robustness. For longer autonomous work, METR's time-horizon research (opens in a new tab) is a reminder that task length and success criteria matter, and that results depend on the agent setup. A public benchmark cannot be translated directly into a promise about a particular business process.

The 2026 AI Index economy chapter (opens in a new tab) finds that measured productivity gains are clearest in structured work where outputs can be monitored, while business agent deployment remains comparatively early. That supports starting with a measurable workflow and increasing authority only after evidence accumulates.

A cheap model can lose its price advantage through review and correction. A more expensive model can become the less expensive process if its outputs are accepted more often. The reverse can also be true: a flagship model used on easy cases may add cost without changing outcomes.

Compare only options that meet the quality and risk threshold, then evaluate total cost and elapsed time.

Build a practical frontier

Plot each tested configuration by complete cost per accepted case and acceptance rate. Remove a configuration if another delivers at least as good quality for no greater total cost. The remaining options form a useful frontier: perhaps one is cheaper, another faster, another more reliable. Apply hard constraints before selecting among them, such as verified citations or mandatory approval before a payment.

The research behind FrugalGPT (opens in a new tab) and RouteLLM (opens in a new tab) supports testing cascades and routing rather than relying on one model. Their reported gains belong to the studied settings; run a local evaluation before expecting them.

Keep a held-out set of routine and difficult cases. Grade blind outputs with a written rubric, include multilingual and damaged inputs, measure routing mistakes, and repeat after model or data changes.

A hierarchy of measures

At model level, measure correctness, format compliance, source use, tokens and latency. These help debugging but do not show that the organization improved. At workflow level, measure cases closed without recontact, time to acceptance, escalation and correction rates. At organization level, measure team capacity, customer experience, revenue or avoided losses under the same quality standards. Improvement at one level does not automatically move the next. A draft may arrive faster while the approval queue still determines completion time.

Separate the objective, hard constraints and diagnostics. The objective might be complete cost per accepted response. A constraint might require a verifiable policy citation or forbid an unapproved payment. Diagnostics include tokens, prompt size and retry rates. Trouble begins when a diagnostic becomes the target: a team minimizes tokens and worsens retrieval, or minimizes escalations by forcing unsupported answers.

Establish causality

Record a baseline before deployment. Then compare similar cases with and without AI, ideally by random assignment or a staged rollout. Observe not only agent handling time but also recontacts, complaints, supervisor work and decision quality later. Generative AI at Work (opens in a new tab) found different effects by worker experience, a warning against one average return for the whole team.

Improvement after a launch can have other causes: seasonality, a new policy, a different case mix or staff training. When comparing models inside one system, hold the case set and acceptance criteria steady. A provider update should trigger re-evaluation; otherwise the refusal style, output format or cost can change silently.

Avoid misleading averages

One benchmark score hides the distribution of errors. Document workflows care about required fields and rare serious omissions. Agent workflows care about permission, reversibility and the state left after an action. Report tail time and cost, plus the severity of the worst recoverable cases. “Most answers look good” says little about whether the remainder is safe.

Automated judges can grade many outputs quickly but may prefer longer confident prose or miss domain-specific mistakes. Calibrate them against specialist-reviewed cases and revisit that calibration as topics change. For retrieval-augmented systems, grade evidence retrieval and fidelity to evidence separately. RAGAS (opens in a new tab) and ARES (opens in a new tab) supply useful evaluation structures.

The frontier moves

Every point on a cost-quality chart represents a complete configuration: model, prompt, data, tools, router and review rules. Change one component and the point must be measured again. A third dimension often matters: a faster accepted result may be worth a higher price. Hard safety constraints still come before optimization.

Finally, tie evaluation to a real decision. A model improvement that does not change acceptance, staff time or risk may have little financial value for that workflow. A small reliability gain on an expensive exception may justify a large token-price difference.

Running example: evaluating the retailer's support workflow

For delivery questions, define acceptance as the correct status from the order system, the applicable current rule and a response that makes no promise the retailer cannot keep. For an address change, add the authority decision: before carrier handoff the system may prepare an action; afterward it must hand the case to a person. For a disputed return, success may be a useful evidence packet for an employee rather than an automatic customer reply. One business process therefore has several distinct success criteria.

Build a historical case set and then observe a fresh flow. Historical cases make model comparisons repeatable but may miss new formats and policy changes. Fresh work reflects reality but requires careful protection of customers and controlled rollout. Early on, compare the AI proposal with employee action in shadow mode. Later, give employees drafts and measure not only speed but edits, rejection reasons and subsequent customer contacts.

Rejection reasons point to architecture. “Status unavailable” suggests integration. “Old policy retrieved” suggests search and versioning. “Promised a refund too confidently” suggests action limits and output structure. “Waited too long for approval” suggests a queue. Trying to solve all four by switching models is usually more expensive and less reliable than fixing the relevant component.

Evaluating rare harm

Serious errors are uncommon, so a small pilot may see none by chance. No observed incident in a small sample is not proof of zero risk. Add deliberately constructed cases: failed authorization, a request for another customer's data, conflicting statuses, outdated instructions and ambiguous consent. Evaluate whether the system can step down to a safe mode when uncertain.

Cost also has a distribution. A typical reply may be cheap; a long chain of repeated tools may delay an employee and exhaust a budget. For an agent, track per-step limits, tail spending and cases ending in human handoff. More authority should require performance on exceptions, not just a good average.

Evaluation as an operating function

Name an owner for the rubric, data set and model-change decision. That owner must be able to accept an update, narrow the automated scope or roll back. Incident review should preserve prompt, model, source and tool-response versions. Without these records, the team will debate impressions instead of identifying causes.

Regular sampling of accepted outputs can expose errors that complaints do not reveal. A customer may not know the cited rule was stale; an employee may silently correct a draft. Evaluation therefore continues after procurement. It becomes part of running the business process, much like ordinary quality control.

Show measurement uncertainty

A test result is an estimate, not a model constant. A small advantage on a limited sample may disappear on new work. Segment results by task type and show the size of each segment. For a rare serious event, a polished overall percentage matters less than constructed stress cases, failure-mechanism review and limits on authority. Zero incidents in a short pilot does not establish safety.

Check agreement between human reviewers. If specialists disagree about what counts as a “good answer,” clarify the rubric and policy before comparing models. Blind grading reduces the influence of provider names. Disputed cases should be recorded: they may reveal a missing business rule rather than a model failure.

In operation, watch for task drift: changes in language, document suppliers, seasonality, policy or customer behavior. Quality can decline even if the model does not change, simply because the input does. Compare the current case mix with the test set, refresh the sample and preserve a smaller stable regression set for version comparisons. After an incident, add a new class of test rather than only copying the exact failed prompt.

When to stop optimizing

Measurement has a cost too. There is little value in endlessly shaving prompt tokens if approval queues or data quality dominate the delay. End an optimization cycle when another experiment is unlikely to change the route decision and critical constraints are met. Restart when price, model, workload or risk changes materially. Evaluation then supports management instead of becoming a factory for charts.

An efficient AI system is a configuration that reliably produces the required outcome for a particular process.

Next: The full cost of an AI workflow.

Sources

Start with the process

Discuss your workflow

Describe one workflow, its inputs, external actions and cost of error. We can identify the smallest level of autonomy that is safe to test.