Citably.

Notebook ·  Ed. 04  ·  September 9, 2026

AI search quoted a passage, not a page.
Only 3 of 14 came from the first paragraph.

Everyone optimizes the page. The engine takes a paragraph. We ran 20 queries through a synthesized web-answer surface, traced each answer back to the source paragraph it leaned on, and measured how deep into the document that paragraph sat. Fourteen queries produced a usable observation. The median passage sat 9.3% into the page body, but only 3 of the 14 were the page’s opening paragraph.

Finding 01

9.3%

median depth of the quoted passage inside its source page.

Where in a page does the passage an AI answer uses actually sit?

Shallow, but not at the very top. Across 14 usable observations measured on 9 September 2026, the median quoted passage sat 9.3% of the way through the source page’s body text, and 10 of the 14 fell inside the first quarter of the document. Seven of those landed in the first tenth. The tail is real though: one passage sat 66.7% down a pricing page and another 83.2% down a comparison post, so depth is a tendency and not a rule.

0-10%710-25%325-50%250-75%175-100%1DEPTH IN SOURCE DOCUMENT (% OF BODY TEXT) · N=14 · 9 SEP 2026
Ten of 14 quoted passages sat in the first quarter of the source document, yet only three were the page’s opening paragraph. Shallow is not the same as first.

That gap matters more than the median does. If the rule were simply “put the answer at the top,” the first bar would be made almost entirely of opening paragraphs. It is not. Most of those shallow passages were the second, third, or fourth block on the page, sitting under a subhead, after an intro the engine ignored.

Finding 02

3 / 14

quoted passages were the page’s opening paragraph.

Does the intro paragraph get the citation?

Usually not. Only 3 of our 14 observations, about 21%, were the first paragraph of the source document. The other 11 were interior passages. The three that did win were all definitional queries where the page opened by defining the term and nothing above it competed. Every commercial query in the set, the pricing and versus and best-of questions, pulled from somewhere in the body instead.

This is the practical read: a strong intro buys you nothing on its own, because the engine is not scoring your page opening against other page openings. It is scoring your paragraph against every other paragraph in the retrieval pool. That is the same asymmetry we found when we counted how often AI answers name their sources inline back in July: the unit the engine works with is smaller than the unit you publish.

“Across 14 traced AI answers, the median quoted passage sat 9.3% into its source page, but only 3 were the opening paragraph. The retrieval unit is the passage, not the page.”

Finding 03

9 / 14

sat directly under a question-shaped heading.

What did the quoted passages have in common?

A question above them. Nine of the 14 quoted passages sat immediately under a heading that was phrased as a question or opened with a question word: “What is a canonical tag?”, “How much does Marketing Hub Professional cost?”, “How ChatGPT decides what to cite”, “Is Airtable better than Notion?”. That held across depth. The passage 83.2% down a comparison post still sat under a question heading. The heading, not the position, is what the two deepest hits had in common with the shallow ones.

The five that did not have a question heading were mostly product roundups and pricing tables, where the heading was a brand name or a plan name. Those pages still got quoted, so a question heading is clearly not a requirement. It is just the single most common thing in the set, and it is the cheapest thing on this list to go and change. Google’s own documentation describes featured snippets as programmatically extracted from a page, which is the same extraction behavior at a different surface.

Worth naming the thing this does not license. Rewriting your copy “for AI” is not the move, and neither is spinning up a page per question. Google’s guidance on AI features in Search is explicit that its AI surfaces run on the same core ranking systems, so a page has to be indexed and snippet-eligible before any of this matters. We made the same point about the file everyone wanted to be a shortcut when we wrote that nothing is reading your llms.txt yet.

The Attribution Series · Part 1

Write the paragraph
you want quoted.

So what do you actually change on the page?

Make every paragraph survive being cut out of the page. If the retrieval unit is the passage, then a paragraph that depends on the two before it to make sense is a paragraph that cannot be quoted. Name the subject instead of writing “it”. Put the answer in the first sentence and the qualification in the second. Give each section a heading a buyer would actually type, and let the passage under it stand alone.

This post is part one of The Attribution Series. If you would rather see the finding applied to your own site than read about it, the Citably audit process scores passages the same way this measurement does.

Method

How we
measured this.

What exactly was measured, and what are the limits?

On 9 September 2026 we ran 20 buyer-intent and explainer queries through a synthesized web-answer surface, the WebSearch tool, which returns a written answer plus the source set it drew on. For each query we fetched every cited source, split the page into block-level paragraphs in document order, and scored each paragraph against the answer text using inverse-document-frequency-weighted token overlap. The highest-scoring paragraph across all of that query’s sources was recorded as the observation, and we logged three things about it: the midpoint of the paragraph as a percentage of the page’s total body text, whether it was the first paragraph, and whether the nearest heading above it was question-shaped.

Six of the 20 queries produced nothing usable and were logged as misses, so N is 14, not 20. In five cases no paragraph in any cited source cleared the 0.25 overlap threshold, which usually meant the answer had synthesized across several sources rather than leaning on one. In the sixth the pages were client-rendered and returned no parseable body text. Because the misses skew toward answers with no single dominant source, treat every number here as a floor on how concentrated passage retrieval is, not a rate.

Other caveats, plainly. This is one surface on one day, and it is a synthesis-plus-citation surface rather than ChatGPT, Perplexity, or Google AI Overviews measured directly. Overlap matching identifies the paragraph most likely to be the source of a claim; it does not prove the engine read that paragraph. Depth is measured against extracted body text, not rendered pixels, so a page with a long boilerplate footer would read as shallower than it looks in a browser. At N equals 14 a single reclassified observation moves a bin, which is why the shape of the distribution is the finding and the exact percentages are not.

Which of your paragraphs can survive being lifted?

Free citability check.
24-hour turnaround.

We score your pages passage by passage, the same way this measurement does, and hand back the three that are closest to quotable.

Filed by Jake Pereira · Founder, Citably · September 9, 2026 · Atlanta GA · ET

The next pass widens the sample past 20 queries and runs the same depth trace against named engines directly, to see whether the question-heading pattern holds when the retrieval stack changes.