All insightsEvaluation

The number nobody audits

9 min read·Manan Jindal·2026-08-07

Somewhere in your finance system there is an invoice for AI work. It has a unit price and a quantity. The unit price was negotiated. The quantity — the number of conversations resolved, tickets closed, invoices matched, claims adjudicated — was produced by the vendor's own software, marking its own work complete.

Nobody checks it. There is no mechanism to check it. And for most of the last two years that did not matter, because AI was a line item in an innovation budget and the amounts were small.

That changed in 2026.

Outcome pricing arrived before outcome measurement

The shift is not coming; it has happened, and it is visible in published price lists. Zendesk bills $1.50 per automated resolution on committed volume, $2.00 pay-as-you-go. HubSpot's Breeze Customer Agent moved to $0.50 per resolved conversation in April. Intercom's Fin offers large accounts a guaranteed 65% resolution rate, backed by a seven-figure payout if it misses. Sierra built its business on outcome pricing from the start.

Enterprise procurement has followed the vendors. Automation minimums, audit-trail requirements, and financial penalties for missed targets now appear in core BPO contract language rather than in an appendix nobody reads.

This is a genuinely good development. Paying for outcomes rather than seats aligns incentives in a way the old business-process outsourcing model never did. But it introduces a dependency the industry has not yet built for: the outcome is now the invoice, which means the measurement of the outcome is now an accounting control.

Measurement built for a dashboard is not measurement built for a contract. Those are different engineering problems with different tolerances, and most teams are still solving the first one.

Deflected is not resolved

The clearest illustration of the gap comes from Gartner's 2026 work on enterprise contact-centre AI. AI deflects more than 45% of customer queries. Roughly 14% reach genuine self-service resolution.

Sit with that for a moment. Of every hundred queries arriving at a support function, about fourteen are actually solved by the AI. About thirty-one are deflected — a human did not touch them — without the customer's problem being resolved. That thirty-one-point band is work that is counted, invoiced, and reported as complete, and is not complete.

The corroborating signal is re-contact. Industry benchmarks put re-contact within 72 hours at 11.3% on AI-resolved tickets against 8.7% on human-resolved ones. The customers come back. They come back to a queue that someone is still paying for. But the original ticket was closed, the resolution was counted, and the dashboard went green before anyone noticed.

Deflection counts conversations a human did not touch. Resolution counts problems that were actually solved. Conflating them is the single most expensive measurement error in enterprise AI today.

Why the reported number drifts

This is rarely fraud. It is almost always the compound effect of three ordinary mechanisms.

Closure is a system event, not a customer outcome. Every support system needs a state transition to end a conversation. When an agent stops responding, or the customer stops replying, something has to write "closed." The system does what it was built to do. It has no way to distinguish a solved problem from an abandoned one, and if the metric is defined on closure, both count.

Judges inherit their author's optimism. Teams that do measure quality increasingly use an LLM to grade the interaction. This is the right instinct. But an unvalidated judge is not a measurement instrument — it is a second opinion from a system with the same blind spots as the first. If nobody has established how often the judge agrees with a human expert on the same sample, the resulting number has an unknown error bar, and unknown error bars trend optimistic when the person building the judge is also the person whose number it is.

Distributions move, evaluation sets do not. A vendor changes an invoice template. A product launch generates a class of question nobody anticipated. A model gets upgraded. The evaluation set was assembled in March and still reports 94%. Half of organisations report having shipped an AI feature that cleared internal evaluations and then failed in front of a customer; a quarter have seen it more than once.

Underneath all three sits a structural fact that no amount of engineering removes: the party reporting the number is the party being paid on it. That is not an accusation. It is the reason the audit profession exists separately from the accounting profession, and it is why the resolution of this problem is not better tooling.

What verification would actually require

If you wanted to establish, defensibly, that a reported outcome rate is real, four things have to be true. Very few programmes currently have any of them.

A validated judge, with its agreement published. Before an automated grader can score production volume, it has to be checked against human expert labels on a held-out sample, and the agreement rate has to be reported alongside every number it produces. A resolution rate without a judge-agreement figure attached is an assertion, not a measurement. This is the single highest-value thing most teams are missing, and it is not expensive to fix.

A sample that represents production, not the happy path. Evaluation sets assembled from clean examples measure the clean path. Verification requires stratified sampling across the actual distribution — including the long tail of malformed inputs, ambiguous intents, and multi-issue conversations where the failures concentrate.

An error taxonomy with costs attached. "94% accurate" is not decision-useful, because the 6% is not homogeneous. Miscoding a low-value invoice and wrongly denying a claim are different events with different consequences. Verification means classifying failures by type and attaching the cost of each, so the confidence threshold can be set against real economics rather than a round number someone liked.

Ground truth that does not come from the system under test. Re-contact rate, downstream escalation, refund and credit issuance, and CSAT on AI-handled interactions are all observable independently of whether the agent marked the ticket closed. Where the closure metric and the independent signal disagree, the independent signal is the one to believe.

Six questions worth asking your vendor

If you are buying AI services priced per outcome, these are answerable in a single call, and the quality of the answers tells you a great deal:

  1. How is "resolved" defined in the contract, and what event in your system sets it? If the answer is a status field, ask what writes to that field.
  2. What is your judge-to-expert agreement rate, on what sample size, measured when? A vendor with a good answer will have it ready. A vendor without one has not measured what they are billing you for.
  3. What is the re-contact rate on interactions you counted as resolved? This is the fastest way to size the gap. If they do not track it, that itself is the finding.
  4. When did you last refresh your evaluation set, and against what distribution? Anything older than a quarter, in a system taking live traffic, is measuring history.
  5. What happens to the invoice when a resolution turns out not to have been one? Most contracts are silent here. The silence is worth pricing.
  6. Who, other than you, has verified any of this? The most informative question, and currently the one most likely to be met with a pause.

The audit function is arriving

Every market that began pricing on a self-reported number eventually grew a third party whose entire product was standing behind that number. Financial statements got auditors. Security posture got SOC 2. Information security management got ISO 27001. In each case the sequence was the same: the number became commercially load-bearing first, and independent verification followed — usually after something went wrong loudly enough to make it non-optional.

Agentic services started pricing this way in 2026. The verification function does not exist yet in any recognisable form. Audit-trail requirements are already entering enterprise contract language, which is the leading edge of the same pattern.

The teams that get ahead of this will not do it by buying another observability tool. They will do it by treating their outcome metric the way a finance team treats a revenue figure: defined precisely, measured by an instrument whose error rate is known, sampled against reality rather than intention, and checked by someone who is not paid on the result.

---

*Pernicia's AgentProof practice builds and validates the measurement layer underneath outcome-priced AI — including judge validation with published agreement rates, stratified evaluation design, and standards-mapped evidence. If you are being billed on a number you cannot independently confirm, [we should talk](/engage?type=agentproof).*

Want to discuss this work?

Start a conversation about how these ideas apply to your context.