The point of the test
You want to know whether an assistant names your business when a customer asks the question that matters. There is no dashboard for that, so you run a controlled observation instead.
It takes about fifteen minutes a month once set up, and it is the most reliable signal available to a small business, for the reasons set out in Lesson 7: measure your AI search visibility.
Step one: build the prompt set
Write five to ten questions a real customer would type. Rules:
- Use their words. “Who can fix a walk-in cooler tonight in Dandenong”, not “commercial refrigeration repair Dandenong”.
- Never name your business. Naming yourself tests whether a record exists, which you already know.
- Cover different intents. One urgent, one price-led, one comparison, one who-is-best, one specific-service.
- Include the location the way customers say it. Suburb name, not postcode, unless your customers use postcodes.
- Write them down once and keep them fixed. Comparability across months is the whole value.
Example set for a commercial refrigeration business:
- Who repairs commercial fridges in Dandenong?
- Emergency coolroom repair near Springvale, open now?
- How much does a commercial fridge callout cost in Melbourne’s south east?
- Best refrigeration technician for a cafe in Noble Park?
- Who services coolrooms for restaurants in Dandenong?
![]()
Step two: run it under fixed conditions
Logged out. Open a private window. Personalisation and conversation history will otherwise show you a flattering version of reality.
Same engines every time. ChatGPT Search, Perplexity and a Google search that triggers an AI Overview covers most customer behaviour. Add others if your customers use them.
Same wording, no follow-ups. If you refine the prompt when the first answer disappoints, you are no longer running the same test.
Note your location. Assistants infer location from your connection. If you test from home one month and a client site the next, the results are not comparable. Record what you used.
Step three: record what you see
One row per prompt per engine, with four columns:
| Column | What to write |
|---|---|
| Named? | Yes or no |
| Who else was named | Competitor names, in order |
| Sources cited | Which domains the answer leaned on |
| Notes | Anything wrong, hedged or surprising |
The sources column is the most useful and the most ignored. It tells you which of your assets the engine can actually see. If a directory listing appears but your website never does, your pages are not being retrieved, and that is a different problem from not being named.
Step four: read it without fooling yourself
Four traps.
Rephrasing until you appear. Feels great, proves nothing. Keep the set fixed.
Reading one good result as a trend. Variance between runs is real. Compare months, not moments.
Treating being mentioned as being recommended. Appearing in a list of six is different from being the named recommendation.
Ignoring wrong facts because you were named. If the answer names you and states an old number, that is a correction job, covered in the guide on what to do when AI search says something wrong.
Step five: turn it into one action
The output of the test is a decision, not a spreadsheet. Each month, pick the single most likely explanation for the biggest gap and act on it.
- Not named anywhere, no sources of yours cited? Start with profile and listing consistency, Lesson 2.
- Your directory listings cited but never your site? Your pages are not quotable yet, Lesson 3.
- Named for general queries but not specific services? You need a page for that service.
- Competitors consistently named with published prices while you say “contact us”? You know the fix.
Then stop measuring and go do the thing. Run the test again next month, at the cadence described in the guide on how often to re-check.
Why logged-out, location-neutral testing matters
Two forces distort results if you let them.
Personalisation. If you are logged in, the assistant has your history, and your history includes your own business. The answer you see may be flattering for reasons that have nothing to do with how it answers for a stranger.
Location inference. Assistants infer where you are from your connection. Testing from your premises, from home and from a client’s office can produce three different answers to the same prompt, and none of them is wrong.
So fix the conditions and record them. A private window, the same three engines, the same wording, and a note of where you tested from. If you want to know how you appear to customers in a specific suburb you serve, say the suburb explicitly in the prompt rather than relying on inference.
Consistency is worth more than realism here. You are not trying to reproduce a customer’s exact experience, you are trying to produce a measurement you can compare against last month.
A worked month-to-month example
A refrigeration business runs five prompts across three engines in January and again in April, after completing month one and two of the 90-day checklist.
January. Named in 2 of 15 combinations. Sources cited: one directory listing, no pages from their own site. Competitors named consistently across all five prompts.
April. Named in 7 of 15. Sources cited now include their service page on two prompts. Hours quoted correctly in every answer where they appeared, which was not true in January.
That is what progress looks like: not a rank, but a dated pattern showing more prompts naming you and better sources being cited. It also shows what to do next, since the prompts where they are still absent point directly at the pages that do not yet answer those questions.
Record the comparison, pick one action, and run it again next month at the cadence described in the guide on how often to re-check.