Wiki/The CPO agenda/
How to evaluate agentic AI vendors for procurement

How to evaluate agentic AI vendors for procurement

The CPO agenda
·
6 min read
·
Updated July 2026
Joshua Kurian
Joshua Kurian
On this page

To evaluate agentic AI vendors for procurement, run five checks that separate agentic substance from agentic labeling: the autonomy check, the context check, the integration posture check, the escalation quality check, and the learning-loop check. Each has a concrete way to run it in a demo or proof of concept, and each tests what the product does with your records rather than what the vendor says. The same five checks apply to every vendor on the shortlist, with no exemption for the one whose page you are reading.

This wiki treats source-to-pay as work AI agents increasingly carry – matching invoices, reading contracts, clearing holds – while people keep the judgment calls. Most vendor-evaluation guides compare feature lists; this page gives a buyer the test protocol for whether the product in front of you can do that work at all.

The autonomy check establishes which rung the product works on

When you evaluate agentic AI vendors, autonomy is the first claim to test, because it is the one most stretched by marketing. Ask the vendor to place the product on the resolution ladder: does it route cases to the right person, suggest a resolution for a person to execute, act under case-by-case human approval, or complete cases end to end within policy? Routing and scripted execution are mature technology, and the difference between them and an agent that investigates is the subject of how agentic AI differs from RPA. The top rung is what autonomous procurement actually means; autonomous exception resolution describes it case by case.

Then verify the claimed rung with one number: the zero-touch count, cases completed with no human event between hold and close. The count is auditable – ERP workflow logs and change documents record every user action on a case, so a reference customer can pull it in an afternoon. If the vendor's headline metric is a recommendation-acceptance rate, you are looking at assistance, whatever the label on the deck says.

The context check is whether it can read your records

An agent resolves cases from evidence, and the evidence lives in your contracts, your purchase order history, and your resolution precedents – so test it on yours. Bring a real contract PDF with an index-adjustment clause and a real failed three-way match, the line-level comparison of invoice, purchase order, and goods receipt that gates payment. In practice, that looks like this: a PO written at $1.24 per pound for 38,000 pounds of resin, an invoice at $1.31, a 5.6% variance against a 2% tolerance (the variance band allowed before review), and a supply agreement whose pricing exhibit resets the price quarterly against a published index. Watch whether the product finds the clause, retrieves the quarter's index value, computes $1.31, and proposes to clear the hold with the calculation attached. A demo environment can be tuned in advance; a scanned exhibit B from your own repository cannot.

The second half of the check is one question: what does the product need connected before it becomes useful? An agent that starts producing evidence from your ERP, contract repository, and shared mailbox was designed for your records. One that needs a data-normalization project first was designed for its own.

The integration posture check asks what you keep if you leave

Establish early whether the product works on top of the existing SAP, Ariba, or Coupa environment or whether its value depends on migrating data and processes into the vendor's platform. The follow-up question does most of the work: what happens to the agent's work product if you leave? If resolutions post in your ERP as ordinary transactions with the rationale attached, the audit trail, the cleared holds, and the resolution history survive the contract. If the record of what was decided and why lives in the vendor's system, switching costs compound every month the agent runs, and the evaluation you are doing now is the last one you will conduct from a position of strength.

Escalation quality shows in thirty seconds

Some cases should reach a person – a supplier claiming a verbal agreement, a variance above sign-off thresholds – and the quality of those handoffs is where products differ most in daily use. Ask to inspect five real escalations, from a reference customer or from your own proof of concept. The difference is visible in thirty seconds. A forwarded symptom ("line 3 price mismatch, please advise") means the analyst starts the investigation from zero. A completed workup – what was checked, what each record showed, the calculation that almost cleared the case, and the single open question that remains – means the analyst makes a decision in minutes. Exception escalation best practices covers what a good handoff contains; in an evaluation you only need to recognize one.

The learning-loop check follows what a correction becomes

Every agent gets cases wrong early, so the question that matters is what happens after a person corrects it. When an analyst overrides the agent – this supplier always bills freight as a separate line, stop flagging it – ask whether that correction becomes precedent applied to future cases with the same pattern, or whether the same mistake recurs until the next model release. Then ask for the correction-rate curve from a real deployment's first months. A falling curve is what learning looks like; a flat one means every case is the agent's first day.

A proof of concept is where you evaluate agentic AI vendors on evidence

Demos reward preparation and a proof of concept rewards substance, so design the POC for a clean read. Four elements do it:

  • A clean baseline. Measure the queue before the agent touches it – current aging, touches per case, escalation share – or the after picture will have nothing to stand against.
  • Metrics agreed up front. Zero-touch share, the collapse in case aging, and the correction rate, in writing before the POC starts, so success is defined before anyone has an incentive to redefine it.
  • A holdout. A comparable slice of the queue worked the old way over the same period, so seasonality and mix shifts do not get credited to the agent.
  • A defined exit. Thresholds that end the POC in either direction, agreed in advance.

The setup is worth it, because this choice sets the operating model your team runs on for years; what agentic AI changes for the CPO covers how far downstream the decision reaches.

Four red flags outweigh a good demo

Four answers should sharply lower your confidence, whatever else went well:

  1. Promises that exception rates go to zero. Exceptions are created upstream, by stale price masters and suppliers who invoice before goods ship, and resolving them faster changes none of that – why exceptions never go to zero works through the causes. A vendor promising an empty queue is describing a product that cannot exist.
  2. Accuracy percentages with nothing behind them. A 98% figure means little without the correction-rate audit that produced it: who counted, over what period, on whose cases.
  3. Demos that run only on the vendor's own data. If the product cannot face the context check on a document you brought, the demo measured rehearsal.
  4. Roadmap answers to present-tense questions. "Which rung does it operate on today" and "show me five escalations" have present-tense answers. A roadmap is an honest answer to a different question.

Fragment builds AI agents that resolve procurement exceptions autonomously inside a company's existing SAP or Ariba environment, and it expects to be measured by this exact protocol – the zero-touch count, the contract-clause test, the escalation inspection, and the correction curve are the numbers Fragment's own deployments run on. See how the workflows run or request a demo, and bring your own failed match.

From Fragment
See exception resolution on your own data
Fragment resolves invoice exceptions autonomously across your existing ERP and documents.
Request demo