Wiki/Invoice exceptions/
Invoice exception rate benchmarks, and how agents change the math

Invoice exception rate benchmarks, and how agents change the math

Invoice exceptions
·
6 min read
·
Updated July 2026
Joshua Kurian
Joshua Kurian
On this page

Invoice exception rate benchmarks are published reference figures for the share of supplier invoices that fail an automated validation check and drop into a hold queue. The most widely cited current figure comes from Ardent Partners' State of ePayables 2025 research, which puts the average invoice exception rate at 18.4%. Used carefully, a benchmark tells you whether your operation sits in normal territory; used carelessly, it aims a project at a number counted differently from yours.

The wiki around this page defines source-to-pay for operations where AI agents carry the routine work – matching invoices, chasing goods receipts, assembling evidence – and people keep the decisions that need judgment. A glossary presents a benchmark as a target to hit. This page treats it as a measurement to interrogate, because the counting rules decide what it can tell a leader, and agents change which benchmark deserves the attention.

Invoice exception rate benchmarks start from a short list of verified numbers

The same Ardent Partners research supplies the surrounding averages: $9.84 to process a single invoice and 8.2 days of invoice processing time. Best-in-Class AP groups, the study's top performance tier, post a 47% lower exception rate, 79% lower processing cost, and 79% faster processing time, and they process more than 1.8 times as many invoices straight-through as their peers – straight-through meaning posted and paid with no human touch.

Two features of that dataset matter before any comparison. First, the spread between the average and Best-in-Class is enormous, which says the gap comes from practice, master data, and process discipline rather than industry fate. Second, the 18.4% average aggregates self-reported figures from companies that define an exception differently, and that definitional spread is wide enough to swallow the entire performance gap.

The same operation can honestly report 5% or 25%

An invoice exception is an invoice that fails a validation check, most often three-way matching – the line-by-line comparison of the invoice against the purchase order and the goods receipt. What counts as a failure is a reporting decision, and it moves the rate more than performance does.

Take a manufacturer running SAP that receives 20,000 invoices a month: 12,000 against purchase orders, and 8,000 non-PO invoices for services, utilities, and legal work that route through approval workflows and never enter the match. In one month, line-level checks flag 3,000 of the PO invoices somewhere – a price mismatch on line 4, a quantity over receipt on line 9. Of those, 2,000 clear on the overnight re-match run once a late goods receipt posts, with no person involved. The remaining 1,000 need manual work. That one month supports four defensible rates:

  1. Every line-level flag, PO invoices only: 3,000 of 12,000, a 25% exception rate.
  2. Every line-level flag, all invoices: 3,000 of 20,000, or 15%.
  3. Manual touches only, PO invoices only: 1,000 of 12,000, or 8.3%.
  4. Manual touches only, all invoices: 1,000 of 20,000, or 5%.

All four are arithmetically honest. The 25% figure counts every failed check at the line level, including the 2,000 holds a re-match cleared automatically. The 5% figure counts only invoices a person touched and spreads them across the full population, including 8,000 invoices that were never eligible to fail a match. When a peer claims 6% and your team reports 22%, the first question is which definition each side used.

Touchless rate and exception rate divide different things

The exception rate and the touchless rate look like mirror images, and they are separate instruments. The exception rate counts invoices that failed a check. The touchless rate counts invoices that traveled from receipt to posting with zero human touches, and an invoice can miss touchless for reasons that never raise an exception: it arrived as a PDF that needed manual keying, or it was a non-PO invoice that waited three days for an approver. In the manufacturer above, the 2,000 auto-cleared holds count as exceptions under the broad definition yet still finish touchless, while thousands of clean non-PO invoices raise no exception and still require a human approval. An exception rate can fall while touchless stays flat, and the reverse. Ardent Partners benchmarks the two separately: the Best-in-Class straight-through advantage of more than 1.8 times is a touchless figure and does not follow mechanically from the group's lower exception rate.

Tolerance settings move the number without changing reality

A tolerance is the variance band a company accepts before a mismatch becomes an exception – price differences up to 2%, say, or quantity differences up to one unit; match tolerances and thresholds covers how the bands get set. Tolerances are also the fastest way to manufacture an improvement. Suppose 380 of the manufacturer's 1,000 manual holds are price variances sitting between 2% and 5%. Widening the price tolerance from 2% to 5% removes them from next month's queue and cuts the reported manual-touch rate from 5% to 3.1%. The overbillings those holds represented still post and still get paid; they have simply stopped being measured. Sometimes that trade is right, since a $14 variance on a $9,000 invoice costs more to investigate than to absorb. For benchmarking, the consequence is blunt: identical operations report different rates under 2% and 5% tolerance bands, and a rate that fell after a tolerance change measured the policy.

Agents move the operative benchmark downstream

Invoice exception rate benchmarks were built for a world where every exception consumed analyst hours, so the arrival rate was the whole story. AI agents that investigate exceptions end to end – pulling the contract clause, recomputing the price against it, checking how the last identical mismatch was resolved, posting the match with the evidence attached – break that link. The exceptions keep arriving, because their causes sit upstream of the queue: stale price masters, suppliers who invoice before goods ship, POs cut after the work started. Why exceptions never go to zero works through those causes. What changes is the cost of an arrival.

Once the arrival rate is understood as an upstream fact, the benchmarks that describe performance move downstream: resolution time, the aging curve of the hold queue, and cost per invoice. Those are the dimensions where the Ardent Partners gaps are widest – 79% lower cost and 79% faster processing for Best-in-Class – and the dimensions an agent-run queue moves, because investigation collapses from days of queue-sitting to minutes of retrieval. A 22% exception rate with a four-hour median resolution beats a 12% rate with a nine-day backlog on every measure that reaches the P&L; the cost of invoice exceptions page traces how those measures compound.

Read the counting rules before acting on the number

A benchmark comparison earns a decision after a few checks. Establish the numerator: line-level flags or manual touches, with auto-cleared re-matches in or out. Establish the denominator: non-PO invoices included or excluded. Check the tolerance policy in force when the rate was recorded, and the period, because quarter-end surges and plant mix swing monthly figures. Then name the decision the number feeds: if the goal is cutting cost per invoice or days to resolve, the exception rate is an input, and the downstream metrics are the ones worth benchmarking. After those checks, the most reliable comparison is still your own trend line, computed under one written definition and broken down by exception type and supplier. The external benchmark says whether you are in normal territory; your own queue says what to do about it.

Fragment builds AI agents that work the resolution side of this math – investigating and clearing flagged invoices inside a company's existing SAP or Ariba environment, with no rip and replace. That moves the benchmarks a tolerance change cannot touch: resolution time, aging, and cost per invoice, while the exception rate keeps reporting honestly on what happens upstream. See how the workflows run or request a demo.

From Fragment
See exception resolution on your own data
Fragment resolves invoice exceptions autonomously across your existing ERP and documents.
Request demo