Reference
A glossary for a discipline three years old
Most of these terms did not exist in 2023 and several are still contested. Where the industry has not settled on a definition, we say which one we use and why.
- Answer density
The proportion of a passage that directly answers the question, versus context and transition.
Rerankers favour passages where most of the text is responsive. Padding a page to reach a word count lowers answer density and measurably reduces the chance of citation, which is why long-form-for-its-own-sake performs badly on AI surfaces.
- Chunking
Splitting a document into smaller passages so a retrieval system can index and select them individually.
Chunking is why page-level quality does not guarantee citability. If your explanation only makes sense across four dependent paragraphs, no single chunk answers the question, and none will be selected.
- Content decay
The gradual loss of traffic to a page that once performed well.
Caused by intent drift, stale facts, competitors publishing more complete coverage, or a structure that AI engines cannot retrieve. Each cause requires a different fix, which is why diagnosis precedes rewriting.
- Generative Engine Optimization (GEO)
Structuring content so language models retrieve, trust and cite it in generated answers.
GEO optimises for the passage rather than the page, and for being quoted rather than ranked. Roughly seventy per cent of the work is rigorous technical and editorial SEO; the remainder is prompt research, source-graph analysis and accuracy monitoring.
- Grounding
Constraining a model's answer to retrieved source documents rather than its trained parameters.
Grounded answers carry citations and can be influenced by publishing. Ungrounded answers reflect training data and change only when a model is retrained, which is why live-retrieval crawler access matters more than training crawler access for most brands.
- Keyword cannibalisation
Two or more of your own pages competing for the same query, splitting the ranking signal.
Detected by clustering your URLs against the queries they rank for. The fix is usually a merge with a redirect, not a rewrite of both pages.
- Passage-level scoring
Grading each retrievable chunk of a page independently for citability.
Distinct from page-level content scoring, which compares a whole document against currently ranking documents. A page can score well at page level and badly at passage level — that gap is the most common reason well-optimised content is never cited.
- Prompt set
A fixed, versioned collection of buyer-intent questions used to measure AI visibility over time.
The AI-surface equivalent of a tracked keyword list. It must be locked and versioned between measurement periods, or trend comparison is meaningless. Good prompt sets come from sales calls and support tickets, not from keyword tools.
- Retrievability
Whether an AI system can reach, render and parse your content at all.
The first of three stages between your page and an AI answer. Failures are mechanical: blocked crawlers, JavaScript-only rendering, interstitials, or non-200 responses. No amount of writing quality compensates for failing here.
- Source graph
The set of third-party domains an AI engine repeatedly cites when answering questions in a category.
The AI-surface analogue of a backlink profile, and usually a much shorter and stranger list. It frequently includes forums, niche trade publications and documentation sites that would never appear on a conventional link target list.
- Topic cluster
A pillar page plus supporting spoke pages, interlinked, covering one topic comprehensively.
Clusters should be built from SERP overlap — comparing which URLs actually rank for each query — rather than from text similarity, which reliably groups unrelated queries that happen to share words.
Want these explained at length? The blog goes deeper.
See what the models say about you.
Run a free visibility baseline across seven AI engines. No card, no call, results in about twenty minutes.
7-day trial · No card required · Cancel anytime