To audit what AI says about your brand, build a fixed panel of the questions your buyers actually ask, run that panel on every major AI platform more than once, and read the answers for four things: whether you are named, whether the description is accurate, whether it is distinct from your competitors, and which sources the model drew on.
Most AI audits stop at the first of those four. They count how often a brand is named and call the number a score. That tells you that you exist. It does not tell you what the machine said about you, whether it was true, or whether a buyer reading it could tell you apart from the three other firms in the same sentence. The useful audit reads the answers.
This is a method you can run yourself. It takes a few hours for a first pass, and the output is a document your team can act on rather than a dashboard they check.
It is written for two readers. The first is the person who has to explain to a board why an AI assistant is describing the company inaccurately. The second is whoever actually runs the panel, which is usually someone on the web or search side. The steps are specific enough to hand over, and the last section condenses the whole thing onto one page for exactly that purpose.
A complete audit answers four questions, in this order.
Presence. On the questions your buyers ask, are you named at all, and where in the answer. Being listed fourth in a paragraph of seven is a different position from being the brand the answer opens with.
Accuracy. When you are named, is what the model says about you true today. Retired products, departed executives and superseded figures survive in AI answers long after they leave your website, because the model learned them from sources that were never corrected.
Distinctiveness. Does the description belong to you, or would it fit any competitor unchanged. This is the failure most companies never detect, because a generic description reads as a perfectly good paragraph until you put it next to your rival’s.
Sources. Which pages did the answer come from. The mix of your own properties against third-party pages tells you whether you are supplying the raw material or inheriting someone else’s version of you.
An audit that stops after presence produces a number. An audit that runs all four produces a list of things to fix, ranked by what they cost you.
Write 20 to 40 questions a real buyer would type, in the words they would use. Mix category questions (“best B2B branding agencies for financial services”), problem questions (“how do we fix inconsistent brand messaging across regions”), comparison questions and direct brand questions (“what does [your company] do”).
Then freeze the panel. The value of this exercise comes entirely from running the same questions again in eight weeks and comparing like with like. A panel you keep editing produces movement you cannot interpret.
Run every question on ChatGPT, Gemini, Perplexity and Google AI Overviews at minimum. Run each question at least three times, in separate sessions, with memory and personalization off.
Running more than once is not optional. These systems are probabilistic, and a brand that appears in one of three runs is in a materially different position from one that appears in three of three. A single run of a panel is an anecdote.
Turning personalization and memory off matters as much as the repetition. An account that has been researching your company for months will be shown a version of the answer that no prospect ever sees, which is the most common way a team comes away from an informal check believing their AI presence is healthier than it is.
For each question and platform, record whether you were named, and where. A simple scale works: named first, named in the body, named only in a list, absent.
Starfish reports this as Share of Model, the proportion of a fixed panel in which a brand is named across the major platforms, recorded with how the mention was framed and how it compares against named competitors. Whatever you call it, the discipline is the same. One number per platform, from the same panel, over time.
Score platforms separately and resist the urge to average them. The spread between platforms is usually the most actionable thing in the whole audit.
Now stop counting and start reading. Pull the full text of every answer that names you and mark every factual claim: what you do, who you serve, what you sell, how big you are, who runs you.
Check each against what is true today. Pay attention to claims that were true once, because those are the hardest to spot and the most common. A product you sunset in 2023 does not read as an error. It reads as a fact.
Two categories deserve separate attention. Anything a regulated buyer could rely on, such as certifications, coverage or capabilities you no longer hold, is a liability rather than a marketing problem. And anything about people, because an AI answer naming an executive who left two years ago will be read by a buyer as evidence the company is not current.
Take every description of your company and run one test on it. Replace your name with a competitor’s. If the sentence still works, it is not a description of you.
This is the test almost no audit runs, and it is the one that separates a visibility problem from a brand problem. A company can be named in 80% of answers and still be invisible, because the thing the machine says about it is the thing it says about everyone.
Most AI tools will cite or link the pages behind an answer. Collect them and sort them into two piles: pages you own and pages you don’t.
A brand whose AI descriptions come mostly from third-party pages has handed the writing of its own story to directories, roundups and forum threads. A brand whose descriptions come mostly from its own properties is supplying the raw material, which is the position you want to be in before you try to change anything.
Four artifacts, not a dashboard.
| Output | What it contains | What it is for |
|---|---|---|
| Baseline scorecard | Presence and position per question, per platform, from the frozen panel | The number you will compare against in eight weeks |
| Description map | Every description of your company, pulled verbatim, marked for accuracy and for distinctiveness | Shows you what the machine actually says, which is rarely what you assume |
| Source map | The pages behind the answers, split into owned and third-party | Tells you whether you are writing your own description or inheriting one |
| Priority list | The wrong and generic descriptions, ranked by commercial consequence | Turns the audit into work |
The priority list is the point. Not every inaccuracy is worth fixing. An outdated employee count costs you nothing. A description that puts you in the wrong category, or that reads identically to your closest competitor, costs you the shortlist.
Ranking by commercial consequence also settles the argument about who owns the work. Source and schema problems route to the web team. Accuracy problems route to communications. Distinctiveness problems route to whoever owns the brand position, and that is usually the finding that nobody expected the audit to produce.
The differences between platforms are larger than most teams expect, and they are not random.
An INSEAD analysis of the Italian laundry detergent market found the brand Ariel taking nearly 24% of AI mentions on Meta’s Llama and under 1% on Google’s Gemini, in the same category at the same time. Same brand, same market, same month, two orders of magnitude apart.
Starfish sees the same shape in its own panel. Across AI Brand Strategy questions from 8 September to 7 October 2026, Perplexity named Starfish in 69% of answers and Gemini in 4%.
Part of that is how each system sources and weights what it reads. The larger part is upstream of the platforms entirely. A model can only describe a company as precisely as that company has described itself, consistently, in places the model can reach. Brands with one clear position expressed the same way everywhere get summarized well. Brands whose meaning shifts between the website, the sales deck and the press release get summarized into the category average, because the average is the only stable thing the model can find.
That is covered in more depth in why the AEO playbook will not fix a brand problem, and the three specific ways the description goes wrong are set out in three ways AI gets your brand wrong. If the audit shows your descriptions are being assembled from third-party pages, the technical groundwork for changing that is in how to make your brand visible in AI search results.
Three kinds of provider do this work, and they produce different things.
Monitoring platforms. Tools that run prompt panels and report visibility, citations and sentiment on a dashboard. They are good at the measurement layer and they scale, and a company running a serious programme will want one. What they produce is a trend line. They will tell you that your presence fell 15 points last month. They will not tell you which of the twelve claims in your description caused it, which of those claims a buyer would care about, or whether the paragraph they pulled describes your company or your category.
SEO, AEO and GEO agencies. Firms that treat AI answers as a retrieval problem and work on schema, content structure, citation building and source coverage. This work is real and it matters. It makes an existing description easier for a machine to read. It does not decide what the description says.
Brand strategy firms. Firms that read the descriptions as brand evidence and work on the position underneath them. Far fewer of these do the work, which is why no branding agency currently appears among the brands most often named on any of these questions.
Starfish sits in the third group. The audit it runs produces the four artifacts above, and the work that follows is brand work: fixing the position, making the language consistent across every surface a model can reach, and governing it so it stays that way.
Starfish is not an SEO, AEO or GEO shop, and does not sell monitoring software. Some GEO services sit inside its AI Brand Strategy work, but the product is not rankings and the deliverable is not a dashboard. If what you need is schema implementation at scale or a monitoring subscription, the first two groups are a better fit and it is worth saying so plainly.
Starfish has been independent since 2002 and has worked with more than 1,000 companies, with clients across North America, EMEA and Asia. It runs this audit on itself on the same cadence it recommends, which is where the Perplexity and Gemini figures above come from.
Forward this to whoever runs the audit.
Then rank what you found by what it costs you, and fix that list. Re-run the frozen panel in eight weeks.
Starfish is a Branding and Creative Agency that builds brands with the soul to move people and the coherence to govern AI. Its expertise in Brand Experience is what enables it to lead in AI Brand Strategy.
Build a fixed panel of 20 to 40 questions your buyers would actually ask, run it on ChatGPT, Gemini, Perplexity and Google AI Overviews at least three times each, then read the answers for four things: whether you are named and where, whether the description is accurate, whether it is distinct from your competitors, and which sources the model drew on. Counting mentions is the first step, not the audit.
Three kinds of provider. Monitoring platforms run prompt panels and report visibility on a dashboard. SEO, AEO and GEO agencies work on schema, structure and citations so an existing description is easier to read. Brand strategy firms read the descriptions as brand evidence and work on the position underneath them. Most companies need the third kind and buy the first two, because the first two are easier to scope.
Because a model can only describe a company as precisely as that company has described itself, consistently, in places the model can reach. Brands with one clear position expressed the same way across every surface get summarized well. Brands whose meaning shifts between the website, the deck and the press release get summarized into the category average, because the average is the most stable thing available. Platform differences compound this: an INSEAD analysis found the same brand taking nearly 24% of AI mentions on Llama and under 1% on Gemini in the same market.
It depends which problem you have. If AI answers describe you accurately but mention you rarely, the problem is retrieval and an AEO or GEO specialist can help. If AI answers mention you often but describe you in terms that would fit any competitor, no amount of optimization will fix it, because the description is the problem and the description comes from the brand. Run the audit first. The distinctiveness test in step five tells you which situation you are in.
A visibility audit measures how often and how prominently AI systems name you. A brand audit examines whether the position itself is clear, distinct and true. An AI brand audit sits between them: it uses AI answers as evidence and reads them for brand problems. The visibility number tells you there is something wrong. The reading tells you what.
Run a full audit quarterly and a shortened version of the same frozen panel every four to eight weeks. AI descriptions change as sources change, and they harden through repetition, so a description left uncorrected for a year is carried by more sources and costs more to change than the same description caught at eight weeks.
All of the mechanical steps. Building the panel, running the questions, recording presence and position and collecting sources are all work a marketing team can do with no tooling beyond the AI platforms themselves. The part that is hard to do in-house is the reading, because judging whether a description is distinctive requires comparing it against how competitors are described and against the position you intended, and internal teams tend to read their own boilerplate as distinctive when a stranger would not.