AgenTomte

July 29, 2026 · 8 min read

Generative engine optimization: how to get cited, not just ranked

By Sahan, co-founder, systems and delivery

Generative engine optimization (GEO) is the practice of structuring content so ChatGPT, Claude, Perplexity, and Google’s AI Overviews quote and cite it in their answers, not just rank it in a list of blue links. You get cited by adding what researchers call “generative readability”: direct answers, named statistics, and quotable sentences that an AI system can lift cleanly into a summary. It has almost nothing to do with the file most people reach for first.

That file is llms.txt, and we will get to why it barely matters. First, the actual mechanics, with the research behind each one.

What is generative engine optimization, exactly?

GEO is content optimization aimed at a machine reader that extracts and quotes rather than a human who scans and clicks. The foundational study, by Aggarwal and colleagues at Princeton and IIT Delhi, published at KDD 2024, tested nine content interventions across thousands of real queries and found that the right ones lifted a source’s visibility in generative-engine answers by up to 40%. The two biggest single levers were adding citations to the text itself and adding direct quotations, not keyword density and not metadata.

That is the whole category shift from classic SEO. Google ranks a page. An AI engine extracts a sentence from a page. Optimizing for extraction is a different job than optimizing for ranking, even though the two overlap more than the industry admits.

GEO versus classic SEO: what actually changes?

The target changes from “rank in position one” to “be the sentence the model quotes,” and that changes which tactics work. Here is the same content problem, scored two ways.

LeverClassic SEOGenerative engine optimization
Primary goalRank in the top 10 organic resultsGet quoted inside the AI-generated answer
Structural winKeyword-optimized headings, internal linkingDirect, self-contained answers near the top of a section
Trust signalBacklinks, domain authorityNamed statistics, cited sources, direct quotations in-text
Who benefits mostIncumbents who already rankChallengers who write the clearest, most citable sentence
Machine-readable filesrobots.txt, XML sitemapNeither llms.txt nor any new file format is required

The Aggarwal KDD 2024 study is the source for that last row on the challenger side too, and it is the most counterintuitive finding in the whole paper: for a source starting at rank five, adding citations to its own text raised its visibility in AI answers by 115.1%, adding quotations raised it 99.7%, and adding statistics raised it 97.9%. Run the same interventions on a source already sitting at rank one, and visibility fell, by 6.0% to 30.3%. The researchers call this the equalizer effect. GEO does not reward the biggest site. It rewards the clearest one, and it can cost the incumbent ground while it does.

Does llms.txt actually get you cited?

No, and the data on this closed in mid-2026. Ahrefs checked the server logs of 137,210 domains in May 2026, of which 28% published an llms.txt file: 97% of those files received zero requests that month. Of the requests that did arrive, AI retrieval bots made up just 1.1%. The largest single category, at 21.7%, was SEO audit tools checking whether the file existed, and Ahrefs notes its own crawlers make up about half of that.

Google said the quiet part out loud in its AI optimization guide, published May 2026 and updated in July 2026: “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search.” Google states plainly that doing so neither harms nor helps visibility. Between one study showing near-zero bot traffic to the file and the largest search engine saying the file does not move the needle, the case for treating llms.txt as a GEO tactic is closed. Write it if you want a tidy manifest for your own records. Do not expect it to earn a citation.

Which crawlers actually decide whether you get cited?

Not the ones most site owners block by habit, and the naming across vendors is genuinely confusing. Each major AI engine runs multiple bots with different jobs, and blocking the wrong one silently kills your citability while leaving you convinced you are still visible.

OpenAI runs three: GPTBot trains models and has nothing to do with citation, ChatGPT-User fires only when a live user’s query triggers a browse action, and OAI-SearchBot is the one that indexes pages for citation in ChatGPT search. Block GPTBot and nothing about your search citability changes. Block OAI-SearchBot and you disappear from ChatGPT’s answers. Anthropic runs an equivalent split: ClaudeBot trains, Claude-User browses live on a user’s behalf, and Claude-SearchBot builds the index Claude cites from; Anthropic’s own documentation states that blocking Claude-User or Claude-SearchBot reduces a site’s visibility inside Claude. Perplexity documents PerplexityBot as the one that must be allowed in for a site to appear in results at all, and separately documents that Perplexity-User, its live-browsing agent, does not honor robots.txt.

The practical move is a robots.txt audit, not a new file: confirm OAI-SearchBot, Claude-SearchBot, and PerplexityBot are explicitly allowed, and stop treating “AI bot” as one category with one on-off switch.

Does one GEO playbook work across every AI engine?

No, because each engine pulls from a different source mix, and a page optimized for one engine’s habits can be invisible to another’s. Profound tracked 680 million citations across AI platforms between August 2024 and June 2025 and found real divergence: ChatGPT’s top cited sources were Wikipedia at 47.9%, Reddit at 11.3%, and Forbes at 6.8%, while Google’s AI Overviews leaned toward Reddit at 21.0% and YouTube at 18.8%, and Perplexity leaned hardest into Reddit at 46.7% of its citations.

That means a strategy built entirely around ranking well on Google will underweight the forum and community content that feeds Perplexity and, to a lesser extent, AI Overviews. If a client’s audience actually researches on Perplexity, ignoring where Perplexity pulls from is a real gap, not a rounding error.

Is getting cited in an AI answer actually worth the traffic?

Partly, and the honest number matters more than the pitch. Seer Interactive tracked 53 brands across 5.47 million queries and 2.43 billion organic impressions from January 2025 to February 2026 and found that being cited in an AI Overview drove 120% more organic clicks per impression than appearing on the same query without a citation. That is the real lift, and it is large.

But read the comparison the other way too. Per one million impressions, queries with no AI Overview at all produced roughly 33,500 clicks; the same queries with an AI Overview present produced about 20,743 clicks when the brand was cited, and about 9,445 when it was not. Getting cited beats not being cited by a wide margin. Both numbers still lose to a search results page with no AI Overview at all. GEO recovers clicks an AI Overview already took away. It does not add clicks a plain search page would not have delivered on its own.

There is a second reason the traffic that does arrive matters more than its volume suggests: Similarweb’s April to May 2026 data put ChatGPT referral traffic at a 7.1% conversion rate, trailing only paid search at 7.8% and beating organic search by a wide margin. Fewer visitors, higher intent.

Where does GEO actually leave you room to compete?

In the citation gap that classic rankings never showed you. BrightEdge’s February 2026 tracking found that only about 17% of AI Overview citations also rank in the organic top 10 for the same query, meaning roughly five of every six citations come from pages that never made page one. AI Overviews themselves are also becoming the default view: they appeared on about 48% of tracked queries in February 2026, up from about 31% a year earlier.

Put those two numbers together with the equalizer effect from the KDD 2024 paper and the shape of the opportunity gets clear. You do not need to outrank an incumbent to get cited next to them. You need a page that states a real number, names its source, and answers the question in the first two sentences, because that is the sentence the model is scanning for, not the domain’s age.

How do you actually build this into a content operation?

The mechanics above only pay off if they run every week, not once. We install this exact discipline as our content engine: a daily article pipeline structured for both Google ranking and AI-engine citation, with a human approval gate on every draft before it publishes. This post is a working example of that pipeline: researched and drafted by an agent, run through a QA gate that checks the sourcing and voice rules above and fails the build if any of them break, then opened as a pull request that only merges once every gate is green. Anything carrying a claim we have not published before stops and waits for a named human sign-off.

The same structure runs on our own B2B group’s site, covered in full in the inbound machine we run for our own catalogue: a 363-product catalogue with daily publishing and structured data on every page, built for the exact citation behavior described above. If you want the fuller case for why most AI initiatives stall before they reach this kind of daily cadence, we wrote separately about the operations gap that kills most AI pilots.

The plain version

Skip llms.txt as a citation tactic; the May 2026 data says it is not read by the bots that matter. Check that OAI-SearchBot, Claude-SearchBot, and PerplexityBot are explicitly allowed in robots.txt, because those three, not the training bots, decide whether you can be cited at all. Write direct, quotable answers with named statistics and sources in the text itself, because that is what the KDD 2024 research says actually moves citation odds, and it moves them furthest for sites that are not already ranking first.

Describe your content operation and you will get a written plan for a GEO-structured content engine within one business day. No meetings, fixed price, and you own the pipeline outright. Start async.

Tell us what you want automated

Describe the work in writing. You get a written reply within one business day: a fixed-price proposal, a scoping question, or an honest referral out.

Start at /start

▸ written reply within one business day · no call scheduled, ever

Doesn't fit a package? Tell us what you need anyway.

Questions? Ask in writing

no chatbot · a human replies

Ask us anything, in writing

A founder replies within one business day. That is the same promise clients get.