The most common advice for getting cited by AI engines sounds a lot like SEO circa 2012: stuff your pages with the right keywords and wait. A Princeton-led study on a 10,000-query benchmark found that keyword stuffing actually performed worse than doing nothing. The tactics that moved: citations (+30–40% visibility), statistics (+30–40%), and quotations (+41%). Two back-to-back experiments, tracking 775 citation events across up to six AI platforms, confirm the pattern and add operational nuance the benchmark can't.
What the Two Experiments Actually Tested
The first experiment ran over several months for an established brand, tracking 15 commercial-intent keywords across ChatGPT, Claude, Gemini, and Perplexity. The strategy combined listicle placements on sources AI models already cited, supported by PR, guest posts, and organic LinkedIn activity. By the end, the brand appeared for roughly 10–12 of the 15 keywords. Peak keyword presence hit 37.01%. ChatGPT led with 148 citations, followed by Claude (96), Gemini (87), and Perplexity (64).
The second experiment was a 30-day cold start for a SaaS link-building agency with zero AI presence at baseline, expanded to six platforms (adding Google AI Mode and Grok). From a standing start, it produced 298 appearances. The platform distribution flipped: Gemini led with 104, Google AI Mode at 95, Claude at 59, ChatGPT at 32, then Grok and Perplexity in single digits.
One original conclusion from the first test didn't survive the second. That alone should give pause to anyone treating a single GEO campaign as a playbook.
Evidence-Forward Content Beat Keyword Targeting
Across both experiments, content that earned citations shared a profile: comparative, evidence-dense, and structured for direct answers. Listicles accounted for 72.4% of citations in the first test and 85.8% of source mentions in the second. PR contributed 24.1% in round one but just 0.2% in round two, suggesting the channel mix is context-dependent, not universal.
Specific findings for ops teams:
- Answer placement: Pages performed better when the answer appeared within the first 100 words. A key-takeaway block near the top outperformed every other on-page change tested.
- Question-based headings: "How is AI SEO different from traditional SEO?" outperformed "AI SEO vs. traditional SEO" in citation frequency.
- FAQ visibility: FAQ content visible by default outperformed content hidden behind expandable sections.
- Comparative framing: Converting a promotional owned listicle into a comparative one (adding competitors by name) produced a 12.25x increase in mentions. Daily visibility jumped from 22 to a peak of 95 in four days.
That last point connects to something the Princeton benchmark also flagged: fluency optimization plus statistics addition outperformed any single strategy by more than 5.5%. Evidence and structure compound. Keywords alone don't.
The Peer-Set Effect Is the Most Replicable Finding
Both experiments surfaced the same pattern from opposite directions. In the first test, a listicle placed the brand alongside recognized experts (Lily Ray, Aleyda Solis). When those names were removed, citation performance declined within days. In the second test, adding recognized competitors to an owned listicle drove the 12.25x spike. Same mechanism, two verticals, two brands.
The implication: AI models appear to evaluate the entities surrounding a mention, not just the mention itself. Being listed alongside credible peers on a credible source may matter more than the raw authority score of the domain. This challenges the "build your own authoritative page" advice that dominates GEO discourse right now.
Source Decay and Measurement Gaps
Roughly half of cited sources stopped appearing within 30 days in the first experiment. One placement fell from 29 mentions to 11 in a single week. Time to first citation in the second experiment ranged from one to 18 days, and the slowest source to be cited was the brand's own listicle. Four placements were live but hadn't been cited when the measurement window closed.
The trade-off you're accepting with a single measurement check: if you audit visibility once, a week after publication, you'll miss both fast decays and slow pickups. For Marketing Ops teams, this means GEO measurement needs a prompt library with a regular cadence, not a one-time audit. Treat it like a reporting channel, not a project.
What Didn't Generalize
Platform distribution shifted completely between experiments. ChatGPT dominated round one; Gemini and Google AI Mode dominated round two. PR was a major citation driver in one context and irrelevant in another. Owned content was the slowest to earn citations but still contributed a 14% foundation.
The honest read: GEO is directional, not deterministic. The Princeton benchmark gives us confidence that evidence-forward tactics outperform keyword tactics at scale. The two experiments give us operational patterns worth testing. Neither gives us a formula you can run once and trust. The content that gets cited is the content that makes it easy for a machine to extract a credible, evidence-backed answer. The models, it turns out, are reading more like editors than crawlers.