Citations Are Receipts, Not Causes: How AI Actually Picks Brands

Β· Β·
Read in: πŸ‡³πŸ‡΄ Norsk
How ai picks brands

Ask most agencies how to get recommended by ChatGPT and you’ll get a checklist. Add FAQ schema. Publish an llms.txt. Write conversationally. Structure your headings for extraction. A dozen versions of that checklist are in circulation and they mostly agree with each other, which is usually a sign that nobody has tested any of it.

Then somebody tested it.

The first controlled studies of AI-search tactics started landing in May 2026, and the results are awkward for the checklist industry. The tactics that everyone agreed on turn out to do very little. Meanwhile the mechanism that does explain which brands get recommended has almost nothing to do with the technical work most teams are billing for.

The study that should have ended the schema debate

Ahrefs ran the experiment properly. They took 1,885 pages, added JSON-LD structured data, and compared them against 4,000 matched control pages over six months, measuring citations across Google AI Overviews, AI Mode and ChatGPT [1].

No meaningful citation uplift. And in AI Overviews specifically, the pages that added schema showed a statistically significant decline.

This gets misread fast, so it’s worth being precise. The finding isn’t “schema is useless” β€” structured data still does real work for rich results, for Merchant Center feeds, for disambiguating your entity. The finding is narrower and more damaging: bolting schema onto existing pages does not buy you AI citations. The causal story everyone was selling doesn’t survive its first rigorous test.

That should change how any other “do X to get cited” claim reads, including the ones in this article. Pedro Dias made the point sharply when he wrote about the confidence gap in the GEO industry: a field where everyone is certain and nobody has run a control group is a field selling opinions [1]. Treat any tactic without a replicated study behind it as a hypothesis you’re funding.

Duane Forrester made an adjacent argument that explains why the checklists were doomed anyway: LLM guidance doesn’t transfer between engines the way SEO guidance did [2]. Classic SEO advice was portable because the engines converged on shared standards β€” sitemaps, robots.txt, canonical tags, schema.org. Those standards got built because it was in everyone’s interest to build them. For LLMs, that shared layer was never built and, Forrester argues, won’t be. Which means a tactic that works on Perplexity has no particular reason to work on AI Mode.

Citations are receipts, not causes

The clearest explanation of what’s actually happening comes from Myriam Jessier’s work on brand depth [3]. Her framing reorders the whole problem.

When ChatGPT recommends a brand and links to a page, most people read the link as the cause: that page earned the recommendation. It didn’t. The model decided which brand to name first, then went looking for something to attach to it. The citation is a receipt issued after the decision, not the reason for it.

That explains a pattern which frustrates a lot of competent marketing teams: traffic logs show AI crawlers hammering the site, the content is clearly being read, and the brand still doesn’t appear in the recommendation. You supplied the raw material. Somebody else got named.

Optimise for citations and you’re optimising the footnote. The interesting question is what determines the sentence above it.

Parametric weight and retrieval survival

Jessier’s model splits visibility into two layers that behave completely differently, and confusing them is where most strategies go wrong [3].

Parametric weight is your brand’s position in the model’s embedding space. It’s built over years, from dense and consistent presence in training data. You cannot ship it this quarter. What you can do is degrade it. If your brand describes itself three different ways across three properties, the vector gets fuzzy and recall suffers. Jessier calls that brand drift. It’s the closest thing to an own goal in this discipline.

Retrieval survival is whether your content makes it through the live pipeline at answer time. This part moves fast, and it’s where the filtering happens. Google AI Mode fans a single question out into 8 to 12 subqueries. ChatGPT disqualifies roughly 83% of candidate sources before it extracts anything. Underneath all of it sits a hard floor: a RAG site-quality score around 0.4, below which you simply aren’t retrieved [3].

One number should reset how you split budget. Roughly 85% of AI brand mentions come from domains you don’t own [3]. Not your blog. Not your product pages. Third-party coverage. Comparisons, forums, trade press. If your entire AI-visibility programme lives on your own website, you’re competing for 15% of the surface area.

Jessier names three entity levers: entity salience (how prominent you are inside a topic cluster), entity coherence (whether you describe yourself the same way everywhere), and inter-entity relationship density (whether you survive the second and third hop when an agent chains queries together). Coherence is the one most teams could fix this month. Almost nobody audits it.

Why models know your brand and still skip it

There’s a gap in the data that is easy to miss. Only 6% to 27% of frequently-mentioned brands are also top-cited [3]. The model knows you exist. It can describe you accurately if you ask directly. It still won’t put you in the shortlist when somebody asks which product to buy.

Recognition and recommendation are different states, and the thing that moves you between them is co-mention density: how often you appear alongside the category, the problem, and your competitors, in sources you don’t control [4]. A backlink with “click here” as anchor text does nothing for this. A trade publication writing “Brand X’s platform integrates natively with Snowflake” does a lot, because it creates the subject-predicate-object relationship the model can actually encode.

This is the part that makes AI visibility uncomfortable for performance marketers. The lever sits closer to PR and analyst relations than to anything inside a campaign management interface β€” it isn’t a setting anyone can log in and change.

Donna Rougeau’s argument on machine readability points at the upstream cause: brands are frequently invisible to AI because they’re missing, thin or badly disambiguated in the knowledge graph [5]. Optimise downstream of a broken entity and you get marginal returns on real effort. Her conclusion, that SEO people need to become the in-house expert on how their brand is represented as machine-readable data, is the most useful reframing of the role to emerge from this shift.

And a caution on where you get your guidance. Michael King’s case is that taking Google’s public AI-search documentation at face value is a mistake, because the docs lag what the systems actually reward [6]. You can see the internal inconsistency in the open: Google’s own answer on whether llms.txt does anything depends on which product team you ask. That reads less like bad faith than like a large company shipping faster than it can document. Either way, reverse-engineering observed behaviour beats reading the manual.

This is a positioning problem, not a technical one

Strip the machine-learning vocabulary away and the shape of this is familiar. Brand positioning has always been about the slot a brand occupies in a buyer’s head β€” which name comes up when a category is mentioned, and how tightly the brand is bound to the problem it solves. That slot gets built slowly, through repeated exposure across sources the buyer trusts, and it can’t be argued into existence on demand.

Parametric weight is that slot. The head in question belongs to a model rather than a person, and it’s a vector space rather than memory, but the mechanic is the same: an accumulated position, built from consistent presence, that no single publishing decision can shortcut.

Which recasts the GEO checklist industry as a well-documented category error. Marketing has a long history of firms trying to fix a perception problem by changing the product β€” a new name, new packaging, a redesign β€” and finding the position unmoved. Adding JSON-LD to 1,885 pages and expecting recommendations to follow belongs in that lineage. It edits the merchandise and waits for the perception to update.

The shortlisting behaviour is the other familiar part. Buyers handle category overload by keeping a very short mental list β€” a handful of names per category, rarely more β€” and ranking within it. An AI recommendation set is precisely that: three to five names returned by a model that demonstrably knows hundreds. It explains the 6% to 27% figure better than anything in the AI-search literature. The brand isn’t missing from the model. It’s known, and below the cut.

That framing also produces two moves the tactical literature doesn’t offer. Displacing an incumbent head-on is the expensive route and rarely works, because the model already holds a settled answer for the head term. The alternatives are to get co-mentioned alongside the incumbent so the association transfers, or to define a narrower category where no settled answer exists yet. A model asked for “best enterprise CRM” has a ready response; “best CRM for Nordic B2B manufacturers with a two-year sales cycle” may have nothing established at all.

None of this is a stretch, because the underlying mechanism was never specific to advertising. It describes any system that screens most of what reaches it and keeps what fits what it already holds. A RAG re-ranker discarding low-entropy content because it adds nothing the model lacks is that same screening behaviour, implemented in code and running at scale.

The engines do not agree with each other

If you’re still hoping for one AI-visibility strategy that covers every surface, look at the portability numbers.

Across an analysis of 118,000 AI responses spanning ChatGPT, Perplexity, Google AI Mode and Claude, only 11% of cited domains appeared on more than one platform [2]. The other 89% were platform-specific.

It gets stranger inside Google. AI Mode and AI Overviews reach similar conclusions about 86% of the time, but cite identical URLs only 13.7% of the time [2]. Two surfaces, one company, the same underlying index. They mostly agree on the answer and mostly disagree about who deserves credit for it.

Meanwhile the link between ranking and citation has been quietly dissolving. In late 2024 around 75% of AI Overview citations came from Google’s top 12 results. By early 2026 only 38% came from the top 10, and one Gemini release replaced 42% of previously-cited domains outright [2]. The strategic split between those surfaces is covered in more detail in AI Mode vs AI Overviews, and the citation divergence is the strongest argument for treating them as separate channels rather than one “AI search” bucket.

So: rank-based proxies for AI visibility are finished. These are independent citation systems that happen to share a crawler. Not a summary layer bolted onto the top 10.

What high-entropy content actually looks like

The one production tactic that follows cleanly from the mechanism is information density. High-entropy content gets retrieved; low-entropy content gets skipped, because a model can generate low-entropy text itself and gains nothing by fetching yours [3].

Low entropy reads like this: “Choosing the right espresso machine depends on your budget, your kitchen space, and whether you prefer manual or automatic. Regular maintenance will extend its life.” Every clause is predictable. A model can produce that sentence without leaving the building.

High entropy reads like this: “Across 500 shots on the Decent DE1PRO, shot-to-shot temperature variance stayed under 0.3Β°C pulling 18g to 36g at 93Β°C.” Named hardware, specific figures, a stated method. None of it can be guessed. That paragraph survives a re-ranker because dropping it would lose information.

The uncomfortable implication for content teams is that volume stops helping. Twenty generic explainers score worse than one piece carrying original numbers, because the deduplication filters treat the twenty as noise. If you can’t say something the model couldn’t have said without you, the page is unlikely to earn retrieval no matter how well it’s structured.

There’s a human-behaviour finding that reinforces this from a different angle. Across 846,000 Google search sessions, users seeing AI Overviews scrolled back up the page nearly twice as often as users who didn’t [7]. They read the summary, then went back to reconsider. Your listing now has to survive a second look, not just a first scan. That rewards specificity over polish.

Preferred Sources: the one lever your readers control

Almost every tactic in this space is at the mercy of the next model release. Preferred Sources is the exception, and it stays underrated.

Google extended Preferred Sources into AI Overviews and AI Mode in May 2026 [8]. Users mark the publications they want to see more of, and those sources get elevated retrieval priority. By the time Search Engine Journal followed up, 345,000 unique sources had been selected, with roughly double the click-through rate to a source once a user had marked it [9].

What makes this different is that the user controls it. It doesn’t reset when Gemini ships a new version and reshuffles 42% of cited domains. If a reader marks you, that preference persists. For a small independent publication it’s the most durable AI-visibility asset going. The play is unglamorous: ask your actual readers to do it.

On measurement, one new tool has actually shipped. Microsoft Clarity now surfaces the grounding queries behind AI citations β€” the sub-questions Copilot generated on the user’s behalf before answering with your content [10]. It’s the first mainstream tool that shows you what the model asked, not just whether you got cited. Google still has no equivalent for AI Mode, which is where most of its AI-summary impressions now live.

What the evidence supports doing

So what survives? Stripping out everything that hasn’t held up under testing, this is what the evidence still supports.

Audit entity coherence first. It’s cheap, it’s entirely within your control, and inconsistency actively degrades parametric weight. Make sure your brand, category and core attributes are described identically across your site, your knowledge graph entry, your social profiles and your press materials.

Move budget from owned content to third-party co-mention. If 85% of AI brand mentions come from domains you don’t own, an owned-media-only strategy is capped at 15% of the surface. That means trade press, comparison content, analyst mentions, podcasts, all with attribute-rich language rather than bare links. Aim that spend at co-mention with the incumbent rather than in isolation: association transfers far more reliably than volume does.

Publish fewer pieces with real numbers in them. One original benchmark beats ten explainers. If you have proprietary data, that’s your retrieval fuel; if you don’t, generating some is a better investment than another round of “what is” posts.

Run a Preferred Sources campaign. Low effort, durable, user-controlled, roughly 2Γ— click-through once marked. There is no reason not to ask.

Find the category you can be first in. If the model already has a settled answer for the head term, competing there is the expensive route. Define the narrower category where no settled answer exists yet and own it. This is among the oldest ideas in positioning, and it survives the transition intact.

Stop reporting AI visibility as one number. With 11% cross-platform citation overlap, a single “AI visibility score” is an average of unrelated systems. Report per surface or don’t report it. The same logic applies to single-score site audits β€” Cloudflare’s Agent Readiness Score bundles a battery of checks into one figure, and the fair criticism is that the checks don’t apply uniformly across site types, so the headline number misleads without the detail underneath it [11].

What the evidence argues against is the thing the industry is still selling: a technical checklist, applied once, expected to produce citations. The one time somebody tested that theory across nearly 6,000 pages, it didn’t hold. Brand depth is slower and much less satisfying to put in a deck. It also appears to be what actually decides the answer.

If you’re building agentic systems on top of this rather than just measuring it, the same discipline applies β€” see building your own PPC OS. The tooling gets easier every month. Knowing which levers survive contact with evidence is the part that doesn’t.

Sources

1. Mt. Stupid Has A Pricing Page β€” Pedro Dias, Search Engine Journal
2. LLM Guidance Doesn’t Transfer The Way SEO Guidance Did β€” Duane Forrester, Search Engine Journal
3. Brand depth determines what AI systems recommend β€” Myriam Jessier, Search Engine Land
4. What co-mentions reveal about the AI recommendation gap β€” Search Engine Land
5. What makes a brand machine-readable in AI search β€” Donna Rougeau, Search Engine Land
6. Google’s AI search guidance is naive and self-serving β€” Michael King, Search Engine Land
7. 846,000 Google searches reveal how AI Overviews change user behaviour β€” Search Engine Journal
8. Google AI Overviews & AI Mode gain preferred sources β€” Barry Schwartz, Search Engine Land
9. Google Preferred Sources hit 345K, expand into AI search β€” Search Engine Journal
10. Microsoft Clarity Now Shows Grounding Queries Behind AI Citations β€” Dan Taylor, Search Engine Journal
11. All You Need To Know About Cloudflare’s Agent Readiness Score β€” Search Engine Journal

Greg Hal
Greg Hal

Performance Marketing Specialist with 14+ years experience. Writing about digital strategies, data analysis and trends in performance marketing.

Related posts