Nine questions that separate a measurement from a dashboard — and our own answers, published next to them.
Every GEO tool and every LLM visibility tool answers the same question — do the assistants mention you — and the differences that matter are all in the method: which engines it really calls, how it decides a mention happened, how many times it asks, and whether you can check any of it against a stored answer. Below are the nine questions worth asking before you pay for one, why each matters, and our own answer to every one of them, including the three that make us look bad. We have not audited anybody else’s tool and this page names none, because a comparison table built out of guesses would fail the same test we are asking you to apply.
| Ask the vendor | Our answer |
|---|---|
| Which engines does it call live, today? “Supports five platforms” can mean five integrations written, not five engines answering. | Two of five. ChatGPT and Claude are live-tested; Gemini, Perplexity and Google AI Overviews have never made a real call. |
| Do you get the raw answers? A percentage you cannot trace to an answer is a number you cannot check, and it is the number you are being asked to pay for. | Yes — one line per call with the full answer text, every link, the model identifier and the timestamp. |
| How does it decide you were mentioned? String matching counts a similarly-named business in another country as you. That inflates the one metric a vendor uses to sell you fixes. | String match, with a hedge flag. On our own first run that flag was the difference between a reported 33.3% and a confirmed 0%; the archive is what caught it. |
| How many times is each question asked? Answers move between runs. One sample per question is a snapshot, not a rate, and the difference is rarely printed. | Three per question in the paid audit; the published sample run used one, and says so on every page that quotes it. |
| Are the scoring weights published? If you cannot see the weights, you cannot tell whether a score moved because your site improved or because the vendor reweighted it. | Yes, all 100 points: the whole rubric, item by item, with the source of every sub-score. |
| Does it show which sources the engine actually cited? Where the assistant goes for information is more actionable than any score, because it is a list you can go and get onto. | Yes — per answer and aggregated; a full example is published. |
| How does it handle Google AI Overviews? There is no official API. Any AI Overview figure comes from a third-party search-results provider or from someone looking. | We do not cover it. Our only AI Overview data is a hand check of 21 keywords, published with its limits. |
| Does it report absolute “AI visibility share”? Every one of these tools samples a non-deterministic system. A share figure implies a census that nobody has. | No. We report counts from named runs on named dates, and comparisons against your own earlier runs. |
| Who is using it? A tool with no customers is a different risk from a tool with a hundred, and you should be told which one you are looking at. | Nobody yet. No customers, no case study, and the sample run on this site is the only end-to-end run there has been. |
Three limits apply to all of them, ours included, and they are structural rather than a matter of engineering effort:
The disclosure we would want from any tool on this results page:
| Assistant | Status | What that is based on |
|---|---|---|
| ChatGPT | live | live-tested 15 Aug 2026, 12 questions, transcripts kept |
| Claude | live | live-tested 15 Aug 2026, 12 questions, transcripts kept |
| Google Gemini | not connected | no calls: our gateway has no Gemini endpoint |
| Perplexity | not connected | no calls: our gateway carries no Perplexity model |
| Google AI Overviews | not connected | no calls: implemented and unit-tested, never run against the live service |
Two of five is a real limitation and it is why the audit is priced where it is. It is also why every page here says which engine a number came from — a figure labelled “AI visibility” with no engine attached is not a measurement of anything.
Do the manual version first; it costs nothing and it tells you whether you have a problem worth paying to measure.
If after that you want the questions asked properly — three samples, two engines, everything archived and turned into forwardable tickets — that is the audit, and it starts with a free sample on your own business.
A tool that measures how generative engines answer questions about you: which assistants mention your business, in what terms, and which sources they cite while doing it. GEO stands for generative engine optimisation, the AI-answer counterpart of SEO.
The raw answers. Every count it reports should trace back to a stored answer you can read, with the engine, the model version and the timestamp attached. Without that you cannot tell a measurement from an estimate.
Treat them as samples of a non-deterministic system. A percentage is useful against your own earlier percentage on the same method; as an absolute claim about your share of AI answers it is not something anybody can substantiate.
Not to find out whether you have a problem. Ask ten discovery questions in the assistant yourself and read the answers. A tool earns its money when you need the same questions asked repeatedly, archived and compared over time.
ChatGPT and Claude, live-tested. Gemini, Perplexity and Google AI Overviews are not connected, and we count our coverage as two of five rather than five.
geo_radar/audit.py: 100 points across five dimensions, reproduced in full on the GEO audit page.