How to Measure AI Workflow Performance Beyond Deflection

Deflection counts conversations that did not reach a person. A production workflow needs measures that show whether the work was completed, sent back, or stopped.

Published:
Updated:

Key terms

Deflection
The share of contacts that do not reach a person. It does not show whether the underlying work was completed correctly.
Completion
The workflow reached the agreed end state in the system of record, not only in the conversation.
Human override
A person changed or stopped the action the AI workforce unit was about to take or had taken.

Deflection is a contact metric

Deflection answers a staffing question: how many contacts never reached a person. It does not answer an operations question: did the order, case, booking, or update finish in the system of record. NIST's measure function asks for metrics tied to the purpose of the system and to the impact on people, not to a single interface statistic.

TechStrata treats a percentage on a slide as unpublished until the measurement period, the source record, and the difference between a result and a projection are written down. This article does not report a customer percentage. It defines the measures a team should be able to read before one is approved for publication.

Read the outcome in the system of record

Completion, cycle time, exceptions, rework, and overrides are visible when the workflow already writes to a system the operator trusts. If the only number available is generated by the model host, the team is measuring the interface. The customer outcomes page uses the same standard: starting condition, workflow, controls, time period, and measured result.

Separate four kinds of statement before a number is used externally. An observed result comes from an agreed record. A projection is a model of a future period. A benchmark assumption is someone else's figure. A product capability is what the system is designed to do, not what a deployment achieved.

Production measure set

TechStrata measurement framework. Use it to decide which numbers are eligible to publish. It is not itself a result.

  1. 01CompletionThe agreed end state exists in the system of record.
  2. 02Cycle timeElapsed time from trigger to that end state, compared with the baseline period.
  3. 03ExceptionsWork that left the automated path, counted by reason.
  4. 04ReworkRecords that had to be corrected after the action was written.
  5. 05OverrideActions a person stopped or changed, counted separately from ordinary approvals.

Related pages

Sources

  1. NIST AI RMF Playbook: Measure
  2. NIST AI Risk Management Framework 1.0