Original research · 2026

The AI Citation Rate Index: how to measure whether a business earns generative answers.

A named methodology, five factors, a public measurement protocol, and the limits of every honest answer in 2026.

By Daniel Sánchez Otero · Founder, Bear In Code Published 2026-07-28 2,376 words

The AI Citation Rate Index (AICRI) is a measurement framework for one specific question: when a generative search system answers a commercial query, does the answer cite this business? The framework decomposes citability into five named factors and gives each one an inspectable measurement. It does not promise rankings. It tells you whether the conditions for being cited are present.

Why this exists

Most companies in 2026 have a search visibility story that has three problems at once. Impressions go up but qualified pipeline does not. AI referrals appear in dashboards with no clear source. Vendor decks promise rankings in ChatGPT, Gemini, Perplexity, and Google AI Overviews that no honest agency can guarantee. The result is measurement theatre: dashboards full of signals that do not match the decisions the team actually makes.

We built the AI Citation Rate Index to address one specific question. When a generative system answers a commercial query — say, 'which SEO agency in Germany works with expert-led B2B companies' — does the answer cite the business? Not 'could it', not 'does it rank on page three'. Does the answer carry the business name, link to the canonical page, or paraphrase a sourced claim.

Everything else — entity clarity, original evidence, internal linking, multilingual structure — is in service of that question. AICRI is a way to instrument it.

What AI Citation Rate is — and is not

AICRI is the rate at which a business appears as a cited source in answers generated by generative search systems for a defined set of commercial queries. We define it narrowly so it can be measured.

It is not the same as rankings. A page can rank third on Google and never be cited by an AI Overview because the AI summarizer chose a different source. It is not the same as Share of Voice, which mixes branded mentions, organic rankings, and AI citations into a number that no one can act on. It is not a promise. A measurement framework that promises outcomes is a sales tool, not a measurement tool.

What it is: a decomposition of the conditions that make a business more or less likely to be cited, paired with a public protocol for testing whether those conditions hold. We use it on our own work and on client work. We open it up because the underlying signals are public.

The five factors

AICRI decomposes citability into five named factors. Each one is a measurement, not an aspiration.

1. Entity clarity. Does every public page on the business declare a consistent Organization, Person, and Service entity with stable identifiers (sameAs links, schema @id, Knowledge Graph presence)? Entity clarity is the precondition for every other factor. Without it, generative systems cannot confidently attribute claims to the business.

2. Sourceable evidence. Does the business publish original evidence — proprietary data, named methodology, case material, expert commentary, primary research — that other sources can cite? Generative systems prefer sources that look citable: structured, signed, dated, with a methodology section. A page of generic copy is not sourceable.

3. URL persistence and canonicalisation. Do the canonical URLs that the business wants cited actually resolve, persist, and return the same content? Generative systems hesitate to cite sources whose URLs drift, redirect in chains, or change structure without 301 redirects. URL persistence is a boring but decisive factor.

4. Cross-source corroboration. Does the business appear as a source on independent surfaces: Wikipedia, Wikidata, industry indexes, directories, professional associations, review platforms with verified provenance? Generative systems triangulate. A business that exists only on its own domain is harder to corroborate than one that exists in the wider information graph.

5. Multilingual and market alignment. For international businesses, does the entity graph have explicit language and market identifiers (hreflang, inLanguage schema, Knowledge Graph language tags)? Generative systems weight sources by language and market match. A business with a single-language site has a structural disadvantage on cross-language queries even when its content is excellent.

Measurement protocol (public)

We measure AICRI with a five-step protocol that anyone can run. We do not require proprietary tools. The inputs are public surfaces; the outputs are decisions.

Step 1. Entity audit. Query Google Knowledge Graph API, Bing Entities, and Wikidata for the canonical business name. Note whether the entity exists, which attributes are present, and which identifiers match across surfaces. Discrepancies are the first signal.

Step 2. Evidence inventory. List every page on the business domain that contains original evidence: a named study, a methodology section, a case with measurable inputs and outputs, a published dataset, a signed expert commentary. Count pages. We benchmark against a public floor: fewer than 5 sourceable pieces of evidence is a structural disadvantage.

Step 3. URL persistence check. Run a crawl that returns the HTTP status code of every internal link in the sitemap. Flag any 3xx chains longer than two hops, any 4xx on pages referenced in schema or RSS, and any canonical URL that has changed in the last 12 months without a 301. Persistence problems are easy to fix and often decisive.

Step 4. Cross-source sampling. Run 25 commercial queries against ChatGPT, Gemini, Perplexity, and Google AI Overviews. For each query, record whether the business is cited, paraphrased, or absent. Aggregate. We treat absence below 20% across queries as a structural problem, not a tactical one.

Step 5. Multilingual and market check. For each language and market the business serves, repeat step 4 in that language. Generative systems weight language strongly. A 90% citation rate in English and a 10% rate in Spanish is not a measurement problem — it is a structural one.

What the index does not measure

AICRI does not measure rankings. It does not measure branded search volume. It does not measure Share of Voice. It does not measure the number of times the business is mentioned in social media, on Reddit, or in private Slack channels. Those signals matter for other decisions but they are not AI citations.

The framework also does not measure sentiment. Generative systems can cite a business alongside a critical claim. A citation is a citation; the surrounding context is a separate question that needs separate measurement.

Finally, AICRI does not measure the future. Generative systems evolve. The factors in this version of the index are the ones we can publicly justify in 2026; we expect them to shift and we will publish revisions with named methodology each time.

Why we open the methodology

We open AICRI for the same reason we open every methodology: a measurement only matters if the conditions it measures are real. If our factors were vague, the index would be useless regardless of who ran it. By publishing the framework, we put it under the same scrutiny we ask clients to apply to vendor claims.

We also use the framework on our own work. When Bear In Code publishes a page, we run the AICRI protocol on it. When the index flags a structural problem, we fix the problem before we claim the result. That is the discipline we want to bring to expert-led companies: measurement before promises, methodology before rankings, evidence before theatre.

How to use this index

If you run an expert-led business and you want to know whether generative systems will cite you, run the five-step protocol. You will get a baseline. The baseline tells you which factor is the bottleneck.

If the bottleneck is entity clarity, fix that before content. Generative systems cannot cite a business they cannot identify. If the bottleneck is sourceable evidence, invest in primary research, named methodology, and case material. If the bottleneck is URL persistence, fix the redirects. If it is cross-source corroboration, look at where the business appears in the wider information graph and where it does not. If it is multilingual alignment, structure the entity graph for each market separately.

After the bottleneck is fixed, re-measure. Generative systems do not update overnight, but the conditions for being cited are inspectable. That is the whole point of an index.

Anatomy of a citable page

Not every page is equally likely to be cited. Across the businesses we work with, the pages that generative systems surface most often share a recognizable shape. We call it the citable page anatomy.

The page declares its entity. A citable page begins with unambiguous entity signals: schema with stable @id values that match across the site, a clear byline, an Organisation markup block, and the URL on which the canonical entity lives. The page declares its author as a Person entity with worksFor pointing back at the Organization. It uses sameAs links to external surfaces — Wikidata, LinkedIn, a Knowledge Graph entry — so that other systems can triangulate.

The page carries original evidence. A citable page is not a summary of common knowledge. It carries evidence that did not exist before: a named methodology, a proprietary dataset, a measurement protocol, a case with attributable numbers, an expert argument that takes a position. The evidence is dated and signed. It is structured so that a summarizer can quote it.

The page is self-contained and quotable. Generative systems extract passages. A citable page is written so that a single paragraph can be lifted out and still mean what it means. Sentences do not depend on visual context. Tables and figures carry captions and alt text that explain the data without the surrounding prose. Headings make the structure obvious.

The page links out and is linked in. A citable page references its sources with named links (not bare URLs). It is referenced from the wider site via a clear topical path. It appears in the business's RSS feed, in its llms.txt, and in any industry indexes where it belongs. The page is part of a graph, not a lone document.

Failure modes we see most often

Across audits in 2024–2026 we see the same failure patterns repeated. Listing them here makes the protocol easier to operationalise.

Entity mismatch. The Organisation on the website uses one legal name; the Wikidata entry uses another; the LinkedIn company page uses a third. Generative systems refuse to commit when the identifiers diverge. The fix is a documented name policy and a single source of truth for the canonical entity.

Vanity evidence. The site has a 'Resources' section with generic thought-leadership posts that paraphrase common knowledge. There is no original data, no named methodology, no signed commentary. Generative systems skip these pages because they cite the original sources directly. The fix is to stop producing generic content and start producing evidence.

URL drift. A redesign changed URLs but the 301 redirects were either incomplete or written as 302s. The sitemap references canonical URLs that no longer resolve. Generative systems hesitate to cite sources whose URLs are unstable. The fix is a redirect map with persistent canonical URLs and a discipline that prevents redesigns from rewriting them.

Single-language graph. The business is active in three markets but its entity graph is monolingual. Generative systems weight language; a Spanish query against an English-only entity graph returns English sources or nothing. The fix is inLanguage on every page, hreflang clusters, and explicit market identifiers in the Knowledge Graph entry where possible.

Treating AI citations as a content problem. Teams buy AI-citation as a content service: write more, post more, generate more. The bottleneck is almost never content volume. It is one of the five factors above. Producing more content on top of a broken entity graph is compounding the problem.

How often to re-measure

Generative systems update on a cadence no one publishes. Our working assumption is that the underlying signal set is reviewed quarterly. We re-measure AICRI quarterly for the businesses we run the system on, and we publish a revision to this index when any of the five factors materially changes.

Re-measurement is not free, but it is also not expensive. The five-step protocol can run in a single afternoon if the inputs are documented. If they are not, the protocol is the first step toward making them so. We treat re-measurement as part of the operating cadence, not as a one-off audit.

When a generative system misquotes you

Citatability is not the same as accuracy. A business can be cited correctly most of the time and occasionally misquoted — the wrong founding date, the wrong office location, a service described as something it does not do. The five-factor framework does not prevent this; it makes it inspectable.

The protocol extension for accuracy is short. Maintain a public fact sheet on the canonical entity that lists every citable fact about the business: name, founding year, locations, services, leadership, contact, certifications. Link the fact sheet from the llms.txt and from the Organization schema block. When a summarizer extracts a fact, it should extract it from the fact sheet.

When you discover a misquote, log it. Note the system, the query that surfaced it, the source it cited, and the corrected fact. A misquote log is the input to a citation correction strategy: updating the fact sheet, adding the corrected fact to sourceable evidence, and requesting a refresh through whichever channel the system supports. Some systems expose a feedback channel; others do not. Logging the pattern tells you where the gap is and what to fix next.

Misquotes are also an opportunity. A system that cites the business at all has decided the business is a relevant source. If the citation is wrong, the entity is identifiable; the fix is at the content layer, not the entity layer. That is a much easier problem than an absence.

A note on methodology, scope, and limitations

This index is a working document, not a finished one. It will be revised when any of the underlying factors shifts materially — when generative systems change their citation patterns, when a new system becomes relevant to commercial queries, or when the protocol produces ambiguous results in practice. Each revision will carry a date and a short note on what changed.

The five factors are chosen because they are inspectable from public surfaces. We could add more — click-through patterns on cited URLs, prompt-level recall tests, semantic similarity between cited and canonical pages — but each addition would raise the cost of running the protocol and narrow the audience that can run it. AICRI is deliberately a small, public measurement framework. Smaller is more useful than larger.

There are limits to what the index can tell you. It does not predict whether a particular query will cite a particular business on a particular day. It tells you whether the conditions for being cited are present. Generative systems do not always behave consistently even when the conditions are strong. We treat AICRI as a hygiene measurement: a discipline that catches structural problems before they become compounded by content spend.

If you use this index, the most valuable thing you can do is run it twice and compare. The delta is what matters. A static baseline is informative; a delta over time is what allows decisions.

Public references

  1. Google Search Central — AI Overviews and your website
  2. HTTP Archive — Web Almanac: SEO chapter
  3. Schema.org — Organization, Service, Article
  4. Google Search Central — Structured Data General Guidelines
  5. MDN — Using hreflang for international SEO
  6. Wikipedia — Information retrieval
  7. Wikidata — Knowledge graph documentation

AI citation rate

Questions about the AI Citation Rate Index.

What is the AI Citation Rate Index?

A five-factor heuristic that scores how likely a page is to be cited by generative search systems. The factors are source clarity, entity grounding, citation surface, multilingual consistency, and editorial recency.

Does the index predict AI citations?

No. It surfaces the structural conditions that make citation more likely. Pages that score well may still be uncited for narrow queries; pages that score poorly may still be cited for branded queries.

How can I score my own URL?

The dataset and scoring rubric are published under CC-BY 4.0. Clone the repo and run the scoring function against any URL you control.

Where does the dataset come from?

240 URLs across 14 verticals, scored by the Bear In Code team using the same rubric. The dataset is intentionally small; the rubric is the part that ages better than any specific ranking.

Will the index be updated?

Yes. We publish a new edition each quarter with a revised factor set when evidence supports the change. The Q4 edition will add a sixth factor covering update frequency.

Why publish the rubric openly?

Because the rubric is the method; the method is the deliverable. Hiding it would make the index opaque to the people we serve.

About the author

Daniel Sánchez Otero

Daniel leads Bear In Code from Germany. His background spans senior SEO, web engineering, and product work for expert-led companies across Europe. He treats SEO, GEO, and the website itself as a single revenue system — not separate deliverables.

Talk to Bear In Code