We ask ChatGPT, Google AI Overviews, Gemini and Perplexity the questions your customers actually type — written natively in their language — and show you how often each one names you, who it names instead, and how confident that answer really is.
Free first measurement · no card · results in minutes
| Platform | Cited | 0% ————— 100% |
|---|---|---|
| Perplexity | 92% | |
| Gemini | 67% | |
| Google AI Overviews | 33% | |
| ChatGPT | 33% |
Citation rates from our published audit. One platform cites a source on nine answers in ten; another on one in three — and carries an interval so wide that a single measurement of it tells you almost nothing.
Cites sources on nearly every answer, and is the most stable of the four.
Grounded answers cite well, but the set of sources moves between runs.
Cites least often of the four, and the interval is correspondingly wide.
Widest interval in the study — a single measurement here tells you very little.
Figures from the multilingual audit, recomputable from the open dataset at 10.5281/zenodo.22491737. The white band on each card is the 95% interval.
A live furniture retailer measured against two named competitors — same questions, same platforms, same intervals. It is a store we operate ourselves, not a customer: with no customers yet, the only measurement we can honestly show you is our own. Nothing here is illustrative — it is one run of the tool, and every figure is in that run's CSV.
of 168 AI answers cited the retailer. The interval is what a 168-answer sample supports — not the single number a competitor tool would print.
Its closest named competitor was cited on 14% of the same answers, interval 7–21%. The two bands do not overlap, so the gap is real and can be stated as fact.
The same brand, the same questions, two languages. Here the intervals do overlap, so the report declines to call it a gap. That restraint is the product.
Everything here exists because the measurements already on the market did not hold up when we tested them.
A rate is never shown without the range it sits in. Where a platform is volatile the range is wide, and you see that at a glance instead of discovering it three reports later.
Movement is only called a change when two measurements' intervals separate. Everything else is labelled within noise. Most tools draw a trend line through variation of exactly this size.
Written natively per market, not translated. German buyers search Nachbau, the Dutch say namaak; neither is what a dictionary returns for "replica".
Every domain cited across the whole answer pool, ranked, with your position in it — so you see who is being recommended in your place, not merely that you were absent.
Worst-performing questions first, broken out by platform and language. A brand strong at home and invisible abroad is the most common finding, and the most actionable.
The prompts, the raw observations and the measurement code are published under an open licence, and your own results export as CSV. Nothing is a black box.
Pick what you sell and the countries that matter. We assemble the buying questions real customers ask in each market, written natively in that language.
ChatGPT, Google AI Overviews, Gemini and Perplexity — every question asked three times, because these systems do not answer identically twice.
Citation rate per platform with its interval, share of voice against every competitor cited, and the questions you lose. Re-measured on your schedule.
Citation rate per platform and per market, each drawn on a 0–100% track with its interval. Wide band, volatile platform. You are never asked to trust a bare number.
Every domain cited across the answer pool, ranked, with your position marked — including when you fall outside the top ten.
Worst-performing prompts first, per platform and market, exportable as CSV so you can work through them.
The right-hand column describes the common pattern across AI-visibility tools we examined while writing the measurement study. Individual products vary, and some do better.
| AnswerRange | Typical AI-visibility tool | |
|---|---|---|
| What is measured | AI search products people actually use | Often raw LLM APIs, which is not the same thing |
| Confidence interval on each figure | Shown on every rate, always | Typically a single percentage |
| Repeat measurements per question | Three, by default | Usually one |
| Change reporting | Withheld unless intervals separate | Trend line drawn through all movement |
| Non-English prompts | Authored natively in six languages | Commonly machine-translated from English |
| Published methodology | Two open-access papers, full datasets | Method generally not disclosed |
| Underlying data | Exportable, CC BY 4.0 | Usually retained in-product |
Generalities are easy. So here is a specific, current product, with only what it publishes about itself — read first-hand from its own pages on 7 September 2026.
| Zutrix — what they publish | What it means for the number you are shown |
|---|---|
| Models tracked | Eight, listed in-product as Claude Sonnet 4.6, GPT-5.5, Gemini 3.1 Pro, Gemini 3.5 Flash, Gemini 2.5 Flash, DeepSeek V4 Pro, Llama 4 Maverick and Mistral Large 3. |
| Three of the eight are Gemini | So 37.5% of a brand's "AI visibility" score is one vendor's models, presented as comprehensive coverage. |
| Google AI Overviews is not among them | Neither is Perplexity. These are the surfaces buyers' own customers use, and the largest of them is absent from a product that describes itself as complete visibility. |
| The headline stat is a fraction of eight | Their "Brand URL Coverage" is the share of the 8 models that linked to you. Three out of eight is 37.5% — a figure whose exact 95% interval runs from 9% to 76%. The interval is roughly 67 points wide, and the single number is what gets printed. |
| One query per model, once a week | Their changelog describes a weekly sync as the measurement, with no repeats. Our audit found a platform contradicting itself on around 40% of repeated identical questions, so week-on-week movement at that sample size is substantially replicate variation. |
Source: https://zutrix.com/pricing and the vendor's public changelog, read on 7 September 2026. We have not seen their output, so nothing above is inferred from it — these are their own published statements and the arithmetic that follows from them. We would apply the same test to ourselves, and do: our method is at /methodology and our data is open.
Every measurement spends real money at four API providers. The plans reflect that rather than an invented seat price.
Prices in euro, excluding VAT. Billing is not yet switched on — every measurement is currently free while the tool is in its first release.
Because a single percentage would overstate what any sample of AI answers can tell you. Ask the same question twice and these platforms often answer differently — in our published audit, one contradicted itself on around 40% of identical repeated questions.
The range is the honest form of the number. A narrow range means the measurement is precise; a wide one means the platform itself is volatile that week. Both are worth knowing, and hiding the second behind a confident-looking figure is how a measurement product misleads you.
It will — once there is a change large enough to distinguish from ordinary variation. Until then movement is labelled within noise.
This is deliberate. Telling a merchant their visibility dropped when it did not is worse than telling them nothing moved, because they will spend money reacting to it.
Mostly in what it refuses to do. The category norm is one question, asked once, reported as a confident percentage with a month-on-month trend. We measured what that method can actually support and published the result: at those sample sizes the smallest detectable change is often larger than the citation rate being measured.
So this asks more questions, repeats each one, and declines to report changes it cannot demonstrate.
There aren't any, because there are no customers yet. This is a first release built on top of two published studies.
We could fill this page with logos and a satisfaction score, as is normal in this category. But the product's whole argument is that measurements should be checkable, and inventing the social proof would refute that argument on the same page that makes it. When there are real users willing to be named, they will appear here.
ChatGPT, Google AI Overviews, Gemini and Perplexity, reached through provider APIs and Google's AI Overview results. These are not identical to the consumer apps — personalisation, session history and app-only retrieval differ. We state that rather than bury it, because it bounds what the number means.
No. Every prompt is authored natively in its market's language. Translation measures the translation: German furniture buyers search Nachbau and Dutch buyers say namaak, neither of which a dictionary returns for "replica". A translated corpus quietly measures the wrong query.
Yes, and that is the point. The measurement study and the multilingual audit are both open access with their full datasets — 4,320 observations, both prompt corpora, the code — under CC BY 4.0. You can reproduce the analysis or disagree with it.
One domain, up to three markets, four AI assistants, every question asked three times. Results in a few minutes, and nothing to cancel.
Already measured something? Your report link is permanent — keep it and re-run it whenever you like.
This tool exists because the measurements already on the market did not hold up when we checked them. Both studies are open access, with the full datasets attached.