Skip to main content
process Free guide

How to Test Whether AI Assistants Recommend You

Build a prompt set that mirrors real customer questions, run it across engines, and record results without fooling yourself.

Updated September 20, 2026 6 min read Part of Measure Your AI Search Visibility
Prompt test sheet listing customer-style questions with result columns per engine

The point of the test

You want to know whether an assistant names your business when a customer asks the question that matters. There is no dashboard for that, so you run a controlled observation instead.

It takes about fifteen minutes a month once set up, and it is the most reliable signal available to a small business, for the reasons set out in Lesson 7: measure your AI search visibility.

Step one: build the prompt set

Write five to ten questions a real customer would type. Rules:

  • Use their words. “Who can fix a walk-in cooler tonight in Dandenong”, not “commercial refrigeration repair Dandenong”.
  • Never name your business. Naming yourself tests whether a record exists, which you already know.
  • Cover different intents. One urgent, one price-led, one comparison, one who-is-best, one specific-service.
  • Include the location the way customers say it. Suburb name, not postcode, unless your customers use postcodes.
  • Write them down once and keep them fixed. Comparability across months is the whole value.

Example set for a commercial refrigeration business:

  1. Who repairs commercial fridges in Dandenong?
  2. Emergency coolroom repair near Springvale, open now?
  3. How much does a commercial fridge callout cost in Melbourne’s south east?
  4. Best refrigeration technician for a cafe in Noble Park?
  5. Who services coolrooms for restaurants in Dandenong?

Results tracking table across three engines over successive months

Step two: run it under fixed conditions

Logged out. Open a private window. Personalisation and conversation history will otherwise show you a flattering version of reality.

Same engines every time. ChatGPT Search, Perplexity and a Google search that triggers an AI Overview covers most customer behaviour. Add others if your customers use them.

Same wording, no follow-ups. If you refine the prompt when the first answer disappoints, you are no longer running the same test.

Note your location. Assistants infer location from your connection. If you test from home one month and a client site the next, the results are not comparable. Record what you used.

Step three: record what you see

One row per prompt per engine, with four columns:

ColumnWhat to write
Named?Yes or no
Who else was namedCompetitor names, in order
Sources citedWhich domains the answer leaned on
NotesAnything wrong, hedged or surprising

The sources column is the most useful and the most ignored. It tells you which of your assets the engine can actually see. If a directory listing appears but your website never does, your pages are not being retrieved, and that is a different problem from not being named.

Step four: read it without fooling yourself

Four traps.

Rephrasing until you appear. Feels great, proves nothing. Keep the set fixed.

Reading one good result as a trend. Variance between runs is real. Compare months, not moments.

Treating being mentioned as being recommended. Appearing in a list of six is different from being the named recommendation.

Ignoring wrong facts because you were named. If the answer names you and states an old number, that is a correction job, covered in the guide on what to do when AI search says something wrong.

Step five: turn it into one action

The output of the test is a decision, not a spreadsheet. Each month, pick the single most likely explanation for the biggest gap and act on it.

  • Not named anywhere, no sources of yours cited? Start with profile and listing consistency, Lesson 2.
  • Your directory listings cited but never your site? Your pages are not quotable yet, Lesson 3.
  • Named for general queries but not specific services? You need a page for that service.
  • Competitors consistently named with published prices while you say “contact us”? You know the fix.

Then stop measuring and go do the thing. Run the test again next month, at the cadence described in the guide on how often to re-check.

Why logged-out, location-neutral testing matters

Two forces distort results if you let them.

Personalisation. If you are logged in, the assistant has your history, and your history includes your own business. The answer you see may be flattering for reasons that have nothing to do with how it answers for a stranger.

Location inference. Assistants infer where you are from your connection. Testing from your premises, from home and from a client’s office can produce three different answers to the same prompt, and none of them is wrong.

So fix the conditions and record them. A private window, the same three engines, the same wording, and a note of where you tested from. If you want to know how you appear to customers in a specific suburb you serve, say the suburb explicitly in the prompt rather than relying on inference.

Consistency is worth more than realism here. You are not trying to reproduce a customer’s exact experience, you are trying to produce a measurement you can compare against last month.

A worked month-to-month example

A refrigeration business runs five prompts across three engines in January and again in April, after completing month one and two of the 90-day checklist.

January. Named in 2 of 15 combinations. Sources cited: one directory listing, no pages from their own site. Competitors named consistently across all five prompts.

April. Named in 7 of 15. Sources cited now include their service page on two prompts. Hours quoted correctly in every answer where they appeared, which was not true in January.

That is what progress looks like: not a rank, but a dated pattern showing more prompts naming you and better sources being cited. It also shows what to do next, since the prompts where they are still absent point directly at the pages that do not yet answer those questions.

Record the comparison, pick one action, and run it again next month at the cadence described in the guide on how often to re-check.

Common questions

Questions readers ask

Does asking the AI about my own business skew results?

Yes, if you name yourself. Asking 'is my business any good' proves only that a record of you exists. Ask the customer's question instead and see whether you appear unprompted.

Should I test while logged in?

No. Personalisation and conversation history skew what you see. Use a logged-out session, and note your location, since assistants infer it from your connection.

How many prompts do I need?

Five to ten real customer questions is enough to see a pattern. More prompts make the routine harder to sustain, and sustaining it is where the value comes from.

What if results differ between runs of the same prompt?

Some variance is normal. That is why you record results rather than react to them, and why you act on patterns across several months rather than a single reading.