All articles
By WEM Editorial Team · Research & price comparison7 min read

How to Audit an AI Shopping Assistant: A Method You Can Run Yourself

Anyone can measure whether an AI assistant quotes real prices. Here is a repeatable method, the scoring rules that keep it honest, and the mistakes that make an audit worthless.

ai-shoppingagentic-commerceprice-accuracymethodologytransparency

Every AI assistant now answers shopping questions, and every one of them will tell you it is helpful. Almost none of them publish accuracy figures, and the few numbers that circulate come from whoever had an interest in the result.

You do not need anyone's permission to check. The method below takes an afternoon, needs nothing but a browser, and produces a number you can defend. It is written so that a journalist, a retailer or a curious shopper can run it and get comparable results.

Decide what you are measuring first

The most common way an audit goes wrong is measuring several things at once and reporting them as one figure. There are at least three separate questions here, and they have different answers.

  1. Did the assistant identify the right product? A price for the wrong item is not a pricing error, it is an identity error, and the fixes are different.
  2. Was the price it quoted correct at the retailer it named, at the moment it answered?
  3. Was the retailer it named actually the cheapest available, or did it miss a better offer?

Score these separately. An assistant that is excellent at the first and poor at the second is a different product from one with the reverse profile, and a blended percentage hides exactly the distinction a reader needs.

Building the question set

Pick products before you pick assistants. Choosing questions after you have seen an answer is how bias enters, and it is invisible in the write-up.

  • Use real products with barcodes, from categories a shopper genuinely asks about: consumer electronics, appliances, beauty, sports equipment.
  • Include a spread of price points. Errors behave differently on a £9 item and a £900 one.
  • Include at least a few products where a cheap accessory shares the product name. This is the identity trap and it is worth measuring deliberately.
  • Write the exact prompt once and reuse it verbatim across every assistant. Varying the wording per assistant makes the comparison meaningless.
  • Aim for enough items that a single odd result cannot move the headline. A few dozen is a study; five is an anecdote.

The part that makes it defensible

Check the retailer's own page at the moment the assistant answers, not afterwards. Prices move. An audit that scores an answer against a page you opened the next morning is measuring the passage of time, not the assistant.

Record, for each answer: the prompt, the assistant's exact words, the price it quoted, the retailer it named, the URL you checked, the price that page actually showed, and the timestamp of both. Screenshots are better than notes. Without this you have opinions.

An accuracy study is only as good as its worst-documented row. Assume you will be asked to defend the one result somebody dislikes.

Scoring rules that stop arguments

Agree these before you score anything, and publish them alongside the result.

  • A tolerance band. "About £50" against £49.99 is the same price; scoring it wrong makes the study look pedantic. A small fixed floor plus a small percentage handles both cheap and expensive items — and state the band you used.
  • Wrong product is scored as an identity failure, never as a price failure, however close the price happened to be.
  • An assistant that declines to answer is not wrong. Count refusals separately. A tool that says "I cannot check that" is behaving better than one that guesses, and a scoring scheme that punishes both equally will reward guessing.
  • Out-of-stock and delisted items get their own bucket. They are a real failure for a shopper but a different one from a wrong number.

That third rule is the one people forget, and it inverts results. If refusals score as errors, the assistant that hallucinates confidently beats the one that admits uncertainty — which is precisely the behaviour nobody wants to encourage.

What to publish

Publish the method before the result if you can, and publish the whole set afterwards rather than the highlights. Two specific commitments make an audit credible.

First, state the sample size next to every percentage, every time. "Eighty per cent accurate" over ten questions is noise wearing a number's clothes.

Second, say in advance what result would count as a failure — including for whoever is publishing. A study whose only possible outcome flatters its author is marketing with footnotes. If you are running this as a business with a product in the category, that pre-commitment is the whole of your credibility.

Reading someone else's study

The same rules run in reverse, and four questions catch most weak work.

  1. How many items, and were they chosen before or after seeing any answers?
  2. Was the retailer page checked at the same moment as the assistant answered, and is the timestamp shown?
  3. How were refusals scored? If they count as errors, the ranking rewards confidence over honesty.
  4. Who paid for it, and what result would have embarrassed them?

None of those questions require expertise in machine learning. They are the ordinary questions you would ask of any survey, and in this category they are rarely asked at all.

The evidence criteria a price has to meet before it can be called verified:

Read the Verified Offer standard

Frequently asked questions

How do I test whether an AI shopping assistant gives accurate prices?

Choose a set of real products with barcodes before you start, write one prompt and reuse it verbatim across assistants, and check the retailer's own page at the same moment the assistant answers. Record the prompt, the answer, the price quoted, the retailer named, the URL checked and both timestamps for every item.

Should an AI assistant refusing to answer count as an error?

No, and counting it as one inverts the result. An assistant that says it cannot check a price is behaving better than one that guesses, so refusals should be counted in a separate bucket. A scoring scheme that punishes refusals and guesses equally rewards confident fabrication.

What tolerance should a price accuracy study use?

A small fixed floor combined with a small percentage, so that honest rounding is not scored as an error while a materially worse deal still is. The important part is deciding the band before scoring and publishing it alongside the result, so readers can see what counted as a match.

How can I tell if a published AI accuracy study is credible?

Check four things: the sample size stated next to every percentage, whether products were chosen before any answers were seen, whether retailer pages were checked at the same moment the assistant answered, and how refusals were scored. Also ask who funded it and what result would have been inconvenient for them.

Get the Sunday deal digest

One email a week: verified price drops and the guides worth reading. Free, unsubscribe anytime.

By subscribing you agree to receive marketing emails. Unsubscribe anytime — see our privacy policy.

Educational content only — not investment, tax, or legal advice. Program rules, rates, and eligibility can change. Refer to the FAQ and terms pages for binding disclosures.

Back to blog