Citably.

Notebook ·  Ed. 05  ·  September 12, 2026

Building a GEO tool taught me
AI search ranks passages, not pages.

I spent months building a scorer that graded pages. It kept telling clients their best page was fine while that page went uncited. The problem was the unit. AI search does not weigh your document against other documents. It weighs one paragraph against every other paragraph in the pool, and my scorer had no idea which of your paragraphs those were.

The Attribution Series · Part 2

The retrieval unit
is smaller than the thing you publish.

What does it mean that AI search ranks passages, not pages?

It means the thing competing for the citation is one paragraph of yours, judged on its own, against one paragraph from somebody else. Classic search retrieves documents and asks which document best matches the query. Generative search retrieves a document and then goes looking inside it for the span that answers the question. Your page is the container. The passage is the candidate.

PAGE-LEVEL RETRIEVALSCORES THE WHOLE DOCUMENTPASSAGE-LEVEL RETRIEVALSCORES ONE SELF-CONTAINED PASSAGE
A strong page can lose to a weaker one that happens to contain a better paragraph. Under passage-level retrieval the rest of your document is context, not a competitor.

I did not arrive at this by reading about it. I arrived at it by watching a scorer be confidently wrong for long enough that I went and traced where the quotes were actually coming from. That trace became the passage-depth measurement in part one of this series, and the headline number is the one that reorganized my thinking: only 3 of 14 quoted passages were the page’s opening paragraph.

What changed

Three habits,
none of them clever.

How do you write a page for passage-level retrieval?

Three changes cover most of it: make every paragraph self-contained, shape your headings like the question a buyer would type, and put the answer in the first sentence under each one. None of these are AI tricks. They are the same edits a good editor would make to a page a human has to skim, which is the part that took me longest to accept.

Self-contained is the one that costs real effort. It means naming the subject in every paragraph instead of writing “it” or “this approach” and trusting the two paragraphs above to carry the referent. When I started editing this way I found that roughly half my body copy could not survive being cut out of the page. Lift it and it stops making sense, because it was leaning on something upstream. A paragraph like that cannot be quoted, no matter how good the argument is.

Question-shaped headings turned out to be less about the model and more about forcing me to admit what each section is for. If I cannot phrase the heading as a question, the section usually does not have a job. Half the headings I wrote before this shift were nouns like “Methodology” or “Considerations”, which is a label, not a promise. Rewriting them as questions killed two sections outright because neither answered anything.

Answer first is the habit I still break. The instinct is to set up the context, then land the point, because that is how you write an argument. Retrieval does not reward the setup. If the first sentence under the heading is throat-clearing, the quotable sentence is buried behind it, and the paragraph reads as hedged rather than authoritative. Say the thing, then qualify it. The qualification is what makes the passage honest, but it goes second.

What I got wrong

I tried to chunk
my way out of it.

Doesn’t this just mean writing shorter, chunkier pages?

No, and that was my first wrong turn. When I learned the retrieval unit was the passage, I assumed the fix was structural: split long pages, chase a target paragraph length, spin up a separate page per question. All three are bad advice. Search systems already handle multiple topics on one page and surface the relevant part, splitting a good page into five thin ones makes each one weaker, and generating a page per query variant is scaled content abuse rather than a strategy.

The honest version is duller. Length is an audience decision. A 3,000-word page made of self-contained paragraphs under real questions is more retrievable than a 600-word page of connected prose, and the reverse is also true. What changed is not how much I write. It is whether each unit of it can stand up alone. Our audit of ten Southeast US B2B SaaS sites found plenty of long, well-researched pages that were invisible for exactly this reason: nothing on them could be lifted.

The second wrong turn was thinking this replaces the fundamentals. It does not. A passage cannot be retrieved from a page that is not indexed, and no amount of paragraph craft outruns a page nothing links to. Passage structure is what decides between you and the other indexed candidate. It is not what gets you into the pool.

Where it leaves me

The scorer got
smaller and more useful.

What does the tool do differently now?

It grades paragraphs and reports the three most liftable ones on a page, plus the ones that fail the lift test and why. A page-level letter grade told a client nothing they could act on. A list that says “this paragraph answers a question nobody asked” and “this one is quotable if you name the subject in sentence one” is an edit list. That is the whole change, and it made the sample audit report about half as long and considerably more useful. The same logic drives how we score entries in the AI Citation Index.

If you take one thing from this: stop asking whether your page is good and start asking which paragraph on it you would be happy to see quoted with your name attached and nothing else around it. If you cannot name one, that is the finding. Every page I have written since has been assembled from paragraphs that pass that test, and the pages read better to humans too, which is the part I did not expect.

Which paragraph on your page is the candidate?

Free citability check.
24-hour turnaround.

We run the lift test on your pages and hand back the three passages closest to quotable, with the edit that gets each one there.

Filed by Jake Pereira · Founder, Citably · September 12, 2026 · Atlanta GA · ET

Part two of The Attribution Series. This one is a field note about my own build, so it cites no outside sources: every number it leans on lives in part one, where the method and the caveats are published in full.