GEO and AI Search: What Actually Gets You Cited by AI (and What Doesn't)
Leggi questo articolo in italiano β
Short answer, up front. There is currently no clean, independent evidence that techniques like llms.txt or JSON-LD structured data measurably increase the chance an AI model cites a site. We built a controlled experiment to test this on our own product, and the primary hypothesis produced no statistically detected effect. What matters most, according to the only source in the field with no commercial stake (Google's own documentation), remains the same old technical fundamentals: crawlability, accessible text content, structured data consistent with what's visible. The rest, for now, is hypothesis β often sold as certainty.
This article explains how we know, what we found, and how you can verify any "GEO" claim yourself before building a strategy on it.
What GEO (Generative Engine Optimization) is
GEO β and its near-synonym AEO, Answer Engine Optimization β is the set of techniques meant to get content cited or recommended by AI answer engines (ChatGPT, Perplexity, Claude, Gemini, Google AI Overviews) instead of, or in addition to, ranking in classic search results. It's the declared heir of SEO for the conversational-assistant era.
The concept isn't the problem. The problem is that much of what's sold as "GEO research" is produced by companies that sell GEO services, with a conflict of interest that's rarely disclosed.
We audited ourselves, and the result was inconvenient
We build Agentabile, a tool that measures how technically readable a site is to an AI agent. We therefore had an obvious interest in proving that the features we measure β llms.txt, structured data β actually increase citations.
We built a controlled experiment to test it, with a discipline rare in GEO marketing: protocol frozen and hash-sealed before collecting data, numbers recomputed from raw data with different code than produced them, human review of the parser on a verifiable sample. The entire study β protocol, code, raw data, review β is public: github.com/AlCap27/audit-black-box-v2-context.
What emerged:
- The primary hypothesis (llms.txt) produced no statistically detected effect. Note the phrasing, because it's where GEO marketing goes wrong in the opposite direction: not detecting an effect is not the same as proving none exists. The intervals were wide; the study lacked power to rule out a small effect. The correct statement is "no effect detected," not "null effect."
- JSON-LD: no effect detected.
- The one positive signal β an advantage for duplicated text content β vanished when we added a control that neutralized preprocessing. This strongly suggests it wasn't a model "preference" but a mechanical effect of the underlying retrieval system (BM25). With caution: the control removed several elements at once, so it points to the direction without definitively isolating the whole effect as an artifact.
An honest and important limit: this is an in-vitro experiment on a synthetic pipeline (lexical BM25 retrieval, a toy corpus, a single generative model, one domain). It says nothing about the behavior of real purchasing agents. We say so first, because it's exactly the over-interpretation we criticize in others.
The method: how to verify any GEO statistic before believing it
This is the reusable part, applicable to anyone β us included. Five checks, ten minutes.
1. Is the "independent study" actually independent?
Search the authors' names together with the company in their affiliation line. Repeatedly, behind an "independent study" cited by a GEO vendor sits another GEO vendor selling the same service. Not proof the data is false; proof it doesn't count as independent corroboration.
2. Is it actually peer-reviewed?
A paper on arXiv, Zenodo or SSRN is a preprint: anyone can post it, it hasn't passed peer review. If the only venue is a preprint repository and no journal or conference accepting it is named, treat it as unreviewed β regardless of how it's cited elsewhere.
3. Do the sources cite each other?
For every "source" a vendor cites, ask what it sells. If the answer is "the same category of service," it's not external corroboration: it's the same market talking to itself. It's called citation laundering.
4. "No effect detected" is not "no effect"
A wide confidence interval is an admission of uncertainty, not proof of zero. Be wary both of those who turn absence of evidence into evidence of effectiveness, and of those who turn it into evidence of ineffectiveness.
5. Is there a source with a different incentive?
For any GEO claim, check whether a platform's official documentation says anything on it. It often has a different, more boring answer, with much less to sell you.
What's reasonable to do today
By level of evidence:
Basic technical hygiene β which Google, who doesn't sell GEO audits, also converges on: crawlable site, text content actually present in the markup (not only client-rendered), structured data consistent with what's visible, clean heading hierarchy, exposed publication dates. This is the only level with independent corroboration.
Plausible but unconfirmed hypotheses β freshness, Q&A format, schema depth: no independent refutation, but no confirmation either. Treat as hypotheses to test, not guaranteed levers.
Contested claims β whether structure "causes" citation regardless of domain authority: the GEO agencies themselves contradict each other. No independent arbiter resolves it.
Why the "definitive" GEO study probably doesn't exist
The question everyone wants answered β "if I structure X this way, do I really get +N% chance of being cited?" β most likely can't be answered cleanly. A controlled experiment gives you control but measures an artificial system, not real agents. The moment you move to real systems for external validity, you lose control: you can't randomize what's already indexed, nor separate the effect of structure from that of domain authority, which almost always dominates. It's a methodological dead end, not an effort problem.
That's why our position isn't "we'll get you cited more." It's: we tell you, verifiably, what your site technically exposes to an agent β including when it matters less than you'd hope.
Frequently asked questions
Does llms.txt actually help you get cited by AI? There's no independent evidence that it does. Our controlled experiment detected no effect, and Google explicitly states that no special files or dedicated markup are needed to appear in its AI features. It doesn't hurt to implement, but don't expect a measurable citation increase on that basis alone.
Does structured data (schema.org / JSON-LD) increase AI citations? As a specific AI requirement, there's no evidence it does on its own. It remains good general technical practice for entity disambiguation and content consistency β do it well because it's technical hygiene, not because it guarantees citations.
What's the difference between GEO, AEO and SEO? SEO optimizes ranking in classic search results. GEO (Generative Engine Optimization) and AEO (Answer Engine Optimization) aim to get cited by AI answer engines. In practice they share most technical fundamentals; the difference is the end goal (ranking vs. citation).
Are the precise percentages in GEO blogs ("+25%", "6x more citations") reliable? Verify before believing. Check who produced the figure, whether it passed peer review, and whether the cited sources are themselves vendors of the same service. Very often they're self-produced numbers backing a commercial offer, not independent measurements.
Does domain authority matter for AI citations? Studies disagree on how much. The most prudent reading is that authority helps you enter the candidate pool, but relevance, accuracy and clarity of the individual page mostly decide the citation. A niche site with the most direct answer can be cited over a more authoritative but vague one.
Do AI crawlers read JavaScript-rendered content? Generally, AI platform crawlers tend to behave like HTTP clients, reading static HTML more than executing JavaScript like a full browser. A site that exposes its main content only via client-side rendering risks being largely invisible. Content that matters should be present in the server-side markup.
What can I concretely do today that makes sense? Focus on technical hygiene confirmed by independent sources: make the site crawlable, put relevant content in static HTML, use structured data consistent with visible content, keep a clean heading hierarchy, and expose dates. Treat everything else as a hypothesis to test, not a guaranteed lever.
How do I know if my site is readable by an AI agent? You can measure its technical readability with an audit tool like Agentabile, which evaluates identity, structured data, and discovery protocols. No technical audit, though, can promise a citation increase: it can tell you what you expose to an agent, not guarantee the agent picks you.
Is this article optimized to be cited by AI? Yes, using only the technical principles confirmed by independent sources (direct answer up front, clean semantic structure, FAQ, dates). If it gets cited, that won't be proof that GEO "works": it'll be an honest test of those basic fundamentals β consistent with our thesis, not a contradiction of it.
The full study β protocol, code, raw data and parser review β is public on GitHub for anyone who wants to re-run it or take it apart.