If you treat generative-engine optimization as a guessing game, these experiments make it measurable. Two structured GEO tests—one multi-month campaign for an existing brand and one 30-day cold start for a SaaS link-building agency—tracked 775 model citation events and reveal which placements and formats actually win LLM attention.
What the experiments measured and the headline results
Each experiment tracked 15 commercial-intent keywords and checked results manually across multiple LLM-driven platforms. The first test (an existing client campaign) monitored ChatGPT, Claude, Gemini and Perplexity over several months. The second was a 30-day cold-start test that added Google AI Mode and Grok.
Key outcomes: the first experiment logged 148 ChatGPT citations, 96 for Claude, 87 for Gemini and 64 for Perplexity. One comprehensive listicle on Indeed SEO generated 190 mentions and outperformed many other placements. The cold-start test produced 298 appearances in 30 days: Gemini 104, Google AI Mode 95, Claude 59, ChatGPT 32, Grok 4 and Perplexity 4.
Two practical patterns stand out: earned third-party listicles drove most model citations, and platform mix changed which sources and queries a model favored.
What actually moved the needle
1) Listicles, PR and guest posts worked as a coordinated system. In the first experiment, listicles typically introduced a claim, press releases amplified it, and guest posts referenced both. Placements that were coordinated across those channels compounded visibility faster than isolated actions.
2) Start with observed citation targets, not generic authority lists. Outreach that prioritized sources the LLMs were already citing produced higher hit rates than outreach based on standard domain-metric rankings.
3) Earned coverage produced the bulk of citations. Across the tests, third-party listicles accounted for the majority of mentions; owned content functioned as a foundation but did not by itself generate the same lift.
4) Authority, relevance and depth beat frequency. About half of sources stopped being cited within 30 days. One comprehensive listicle consistently outperformed several shallow pieces combined, and placements on higher-authority, relevant outlets showed greater persistence.
5) Peer set and framing matter. Being listed alongside recognized experts increased citation rates; removing those peers reduced performance within days. Similarly, reframing an owned listicle from promotional to comparative (explicitly naming peers) increased its mentions dramatically.
6) Exact-match coverage and answer placement still matter. Pages that placed the key answer within the first 100 words performed better. Exact-phrase assets for commercial queries tended to outperform broader semantic coverage when intent was commercial.
7) Format details influence outcomes. Visible FAQ content outperformed hidden accordions, question-form headings beat neutral headings, and visible publication dates or recent references correlated with stronger citation performance.
8) Platform differences change priorities. In the cold-start test, Gemini and Google AI Mode together accounted for roughly two-thirds of appearances; in the first test ChatGPT led. That means which models you prioritize should depend on where commercial intent concentrates for your vertical.
Concrete steps marketers can take this week
– Run a baseline: query each target keyword across the LLMs you care about and log every source cited. Use that ranked list to prioritize outreach rather than relying solely on generic authority metrics.
– Target a few high-leverage placements: prioritize sites that already surface in model citations. A small number of sources produced around three-quarters of visibility in the cold-start test.
– Make owned content comparative and front-load the answer: for commercial queries, include peer names, place the core answer in the first 100 words, and use question-style headings and visible FAQ sections.
– Coordinate earned channels: amplify listicles with PR and guest posts to create reinforcing signals that the models repeatedly surface.
– Measure citations repeatedly: check multiple times after publication—time-to-citation ranged from 1 to 18 days—so a single check can miss later pickups.
– Track both citations and referral impact: high citation counts do not always drive traffic. Use both citation tallies and referral/session data to decide where to double down.
What next to watch
The experiments clarify immediate tactics, but they also raise new measurement questions: will platform retrieval behavior stabilize around a consistent peer set, can sustained PR campaigns lengthen citation life, and will comparative third-party listicles dominate visibility across more verticals? Practically, watch which platforms drive commercial intent for your keywords and validate placements by measuring both citations and downstream referrals.
For now, treat earned placements on already-cited sources and the first 100 words of an answer as two of your highest-leverage levers when optimizing for generative AI visibility.