◄ ALL POSTS
GEO / AI SEO

Snake oil for AI search

NW Nils Weiser Jul 8, 2026 12 MIN

GEO, "Generative Engine Optimization", is the supposed new discipline for getting cited in ChatGPT, Perplexity, and Google AI Overviews, and it's currently being sold as a paid service. The short version: the traffic shift is real. Almost everything sold as a "GEO strategy" to counter it is not.

The honest answer up front: there is no GEO playbook that does more than two things that existed long before the word GEO: writing substantial, well-sourced content and doing solid technical SEO. Everything beyond that is either crawler plumbing, long-debunked tactics, or simply invented numbers.

// Short on time? Jump to the overview: evidence vs. snake oil

The traffic shift is real, that's the true core

Let's start with what's true. The shift away from classic blue-links search toward synthesized AI answers is measurable and large:

−25 %
classic search volume by 2026 (Gartner forecast)
−58 %
clicks on position 1 with AI Overviews (Ahrefs)
~2.5 bn
prompts/day via ChatGPT

Ahrefs measured that a displayed AI Overview cuts the click rate on the top organic result by 58% (end of 2025), versus 34.5% just eight months earlier. The curve points steeply upward. Anyone optimizing only for position 1 today is optimizing for a shrinking window. So far, the alarm is justified. The mistake starts only afterward: with the assumption that this new window can be "optimized" with the same tools as a SERP.

Why GEO has no measurable substrate

Classic SEO always had an observable output: a results page with positions you could track. Two people, same query, same location produce practically the same SERP. That's the foundation on which ranking tracking makes any sense at all.

LLM outputs don't have that substrate. The answer depends on sampling temperature, model version, exact prompt wording, user context, retrieval strategy, and a constantly changing index. Run the same query twice and you get different citations. There is no "citation rank" you could pin down, only a probability distribution.

The numbers here are sobering. One analysis put the month-to-month variance in AI citations at 40–60%, with no content change at all. ZipTie found that 89% of citations differ between ChatGPT and Perplexity, and only 18% of brands show up simultaneously across all three major AI platforms. So you're not optimizing for a position, but for a flickering noise.

What the Princeton study actually shows

The academic origin of the term is the work by Aggarwal et al., "GEO: Generative Engine Optimization" (Princeton/Georgia Tech/AI2/IIT Delhi, published at ACM KDD 2024, DOI 10.1145/3637528.3671900). It's the most-cited foundation of every GEO sales page, and it's almost consistently misrepresented in the process.

What the study did: a benchmark of 10,000 queries against a system that reconstructed Bing Chat. What it reports:

+41 %
citation probability from statistics
+28 %
from expert quotes
30–40 %
from third-party citations

Sounds like proof that "GEO works". Two caveats that appear on no sales page:

First: these are benchmark artifacts from a controlled environment that doesn't exist in production. The reconstructed "generative engine" is deterministic enough to produce measurable lifts at all. That is exactly not the case with real ChatGPT (see above).

Second, and this is the fun part: SandboxSEO pointed out that the three "winner" techniques (statistics, quotes, sources) all added content, while the six ineffective techniques merely reworded existing text. So the lift might simply be content density and not "optimization magic". Put differently: the study mainly proves that more substantial, sourced content helps, exactly the thing good content has anyway.

Fittingly, the Semrush analysis of 230,000 prompts across 3 LLMs and 13 weeks: domains that rank well in classic search are also cited more often by AI. ZipTie's caveat: domain authority alone explains only about 18% of citation variance. It helps, but it's not a lever, it's a side effect of good substance.

Dashboards that weigh smoke

Spending on "AI visibility tracking" now runs beyond 100 million dollars a year, spread across hundreds of dashboards that promise you a "ranking position in AI". The research on this is unambiguous: that position doesn't exist.

In January 2026, Rand Fishkin and Patrick O'Donnell (Gumshoe.ai) recruited around 600 volunteers and ran 12 brand-recommendation prompts a total of 2,961 times through ChatGPT, Claude, and Google's AI (SparkToro study). The result: ask an AI for brand recommendations 100 times, and the chance of getting the same list twice is below 1:100; the same list in the same order closer to 1:1,000. Fishkin's verdict on any tool that outputs a "ranking position in AI": "full of baloney".

The academic side confirms it. An arXiv investigation, "Don't Measure Once", found a source overlap (Jaccard) of only 0.32–0.43 for an identical prompt within minutes. These values show that internal stochasticity alone explains most of the instability. A second paper, "Quantifying Uncertainty in AI Visibility", puts it bluntly: citation metrics of generative engines are random variables, not fixed values. A single measurement carries an unquantified uncertainty large enough to invalidate common conclusions. Ahrefs' number fits the picture: Google's AI Mode and AI Overviews cite 87% different sources for the same query.

The one stable signal

Fishkin found exactly one thing that stayed stable while the order kept getting reshuffled: the "consideration set", the pool of brands the model draws from in the first place. Membership in that pool can be estimated, but only slowly, over 60–100 runs. That's the only valid signal. Measure membership, not rank.

The "wrong window" problem

And most tools even take that measurement in the wrong place: they query cheap vendor APIs. But GPT-via-API ≠ ChatGPT. The consumer product has memory, custom instructions, history, account context, location, and its own system prompt. Over 90% of weekly ChatGPT users are on the free tier, which triggers fewer searches and fewer citations than the paid tiers the tools sample. On top of that, prompt-volume figures are mostly made up, either modeled from tiny clickstream panels or back-calculated from Google keywords. Steve Toth and William Alvarez called Profound's prompt volumes a "scam"; Conductor described any AI prompt-volume metric as fundamentally misleading.

What the evidence supports and what's snake oil

Here's the split that sums up the entire article. On the left, what there's solid evidence for; on the right, what gets sold but was never cleanly proven.

What the evidence supports Snake oil
Allow AI retrieval crawlers in robots.txt (OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot) llms.txt as a GEO tactic: SE Ranking found no correlation across 300,000 domains; Google doesn't support it (Illyes)
Render content server-side (crawlers don't reliably execute JS) Blocking all AI bots wholesale (Rutgers/Wharton, Dec. 2025: −23% traffic vs. peers that allowed crawling)
Structure for extraction, clear headings, answer up top (AI Overviews cite 55% from the first 30% of the content) Citation guarantees from agencies and tools
Sourced statistics (+41%), expert quotes (+28%), third-party sources inline (30–40%, up to 115% on weaker pages) Schema as a universal lever (Search Atlas, Dec. 2024: no correlation across OpenAI, Gemini, Perplexity)
Original research / proprietary data (ZipTie: high data density yields 4.31× more citations per URL than directory listings) Vendor lift statistics without methodology
Maintain domain authority via classic SEO (Semrush; but explains only ~18% of the variance) "AI visibility" tools with precise citation-share rankings
Schema markup, confirmed only for Bing Copilot (Microsoft/Canel, SMX Munich 3/2025) and probably Google AI Overviews (Search Liaison, 4/2025) Tracking rank instead of membership (see Fishkin)

On the llms.txt debate, Kai Spriestersbach delivered the aptest image: it's as if a restaurant had a menu noting that it reads other restaurants' menus before cooking. John Mueller compared it to the discontinued meta keywords tag; Semrush tested one live on Search Engine Land: no effect.

What this means for your website

No panic-buying a GEO retainer required. The evidence-based approach is cheap, mostly one-time, and overlaps almost entirely with good technical SEO and good content. Concretely:

1 · Do the Bing plumbing

ChatGPT's live search runs mostly on the Bing index (Seer Interactive: 87% of ChatGPT citations matched Bing top results, only 56% Google). So: verify Bing Webmaster Tools, wire up IndexNow, allow OAI-SearchBot in robots.txt. Bing's free AI performance report (since Feb. 2026) shows you real data instead of vendor estimates.

2 · Serve static HTML

The parse rate collapses from 94% for static HTML to 23% for client-rendered JS. A pure React/Vue shell looks like an empty page to an AI crawler. (This site here is static for exactly that reason.)

3 · Fix wrong information at the source

The Tow Center/Columbia found that AI search couldn't retrieve the correct citation information in over 60% of 1,600 queries. Before you chase "visibility": make sure that what's retrievable about you is even correct in the first place.

4 · Measure membership, not rank

If you measure at all: 60–100 runs per prompt, and look at whether your brand shows up in the consideration set, not in which "spot". Everything else is noise with a pretty axis label.

5 · Earn third-party mentions

SE Ranking: pages with over 32,000 referring domains are cited 3.5× more often than those with under 200. Earned media beats any on-page trickery. As an aside: the machine's "fan-out" queries are what matter: ALM Corp found that 95% of these machine sub-questions have almost no human search volume.

How big is the prize anyway?

Finally, the order of magnitude, so nobody misallocates the budget. According to Cloudflare Radar, the combined chatbot share in May 2026 was 0.29% of search referrals, versus 87.6% for Google. The traffic shift is real, but today it's still small.

The nuance that makes it interesting anyway: ChatGPT referrals convert at 7.1% according to Similarweb, only paid search is higher. And around 70% of AI visits arrive without a referrer and get wrongly logged as direct traffic. So the channel is small, but high-value and systematically underestimated. That's exactly why the cheap, evidence-based approach is worth it, but the expensive GEO retainer is not.


Sources:

arXiv 2311.09735 (KDD '24) arXiv 2603.08924 arXiv 2604.07585 SparkToro / Gumshoe betterthangood.xyz (Iain)

Are you building AI systems where whether and how you're found by generative engines matters, or do you want to know what actually holds up in a GEO offer? Let's talk. I build agents that do real work, and I'll tell you honestly where substance ends and snake oil begins.

NW
Nils WeiserAI Agent Specialist · Bodenseeraum
WORK WITH ME ▸

More field notes

ALL POSTS ▸