How a generated answer is built
A generative engine reads several pages and answers in its own sentences, with the sources inside the answer. An earlier web engine handed back a list [1] [2]. The idea underneath is retrieval-augmented generation (RAG): a language model is paired with a search index, so it can draw on documents it was never trained on. In the original 2020 work, this produced more specific, diverse and factual language than a model working from memory alone [4].
Retrieval also brings a new weakness. Models use long inputs unevenly. Performance is often highest when the relevant passage sits at the start or end of the input, and drops when it sits in the middle [5]. A page doesn’t control where an engine places it. It can control whether the answer to the question sits in one clear, self-contained passage.
What the 2024 GEO study measured
Aggarwal, Murahari and their colleagues called this craft generative engine optimization (GEO) [1]. Visibility in their study is not a rank. A source cited early, and woven through more of the sentences, occupies more of what a person actually reads [1]. They measured this with a position-adjusted word count and a judged “subjective impression” score. On the full benchmark, the stronger edits raised visibility by as much as 40 percent [1].
The edits that helped were small enough to miss:
- Adding a citation from a reliable source, a credible quotation or a relevant statistic raised the position-weighted word count by about 30 to 40 percent, and the judged impression by about 15 to 30 percent [1].
- On Perplexity, a public engine, quotations raised the position-weighted count by 22 percent, and statistics raised the judged impression by up to 37 percent [1].
- Keyword stuffing did little or nothing on the authors’ engine. On Perplexity it lowered the position-weighted count by about 10 percent, even while the subjective score rose [1].
- Fluency paired with statistics beat either change alone by a little over 5 percent, in a smaller follow-up of 200 examples [1].
Subject mattered too. Citations suited factual statements, and statistics suited law, debate and opinion [1].
Why tone is the weaker bet
The authors expected a persuasive, authoritative tone might help, because models are built to follow instructions. They found no significant improvement across queries as a whole [1]. Tone fit debate and history better than other subjects [1].
A separate 2024 study reached a related result from another angle. Wan, Wallace and Klein paired contested questions with real web pages that differed in facts, argument style and answer [6]. Models relied heavily on how relevant a page was to the query. They largely ignored stylistic features that people find important, such as whether a text contains scientific references or is written in a neutral tone [6]. That study asked which answer a model accepts, not how much of a page it shows. So it doesn’t overturn the GEO result on citations. It does say that relevance comes first, and that style on its own is a thin lever.
Who gains, and how far the result travels
The finding we return to is about who moves. For the citation method, a page ranked fifth in ordinary search gained about 115 percent in visibility. The page already ranked first gave up about 30 percent [1]. The authors’ reading is that a generated answer depends on the text it was shown, so the links a smaller site has not yet earned need not decide all of its chances [1].
The same paper says engines will change, and useful edits may have to change with them [1]. A 2025 benchmark, C-SEO Bench, tested published editing methods across two tasks and six domains, including cases where many sites edit at once [7]. Most methods were largely ineffective and often lowered a document’s ranking. Strategies that improved a source’s ranking in the model’s context were more effective. Gains also shrank as more sites adopted the same method [7]. We read this as a limit on recipes, not on evidence. A citation that helps a reader check a claim is worth adding even when the visibility gain is uncertain.
Citations in the answer are not always sound
Across four generative search engines audited in 2023, only about 51.5 percent of generated sentences were fully supported by their citations. Only about 74.5 percent of citations actually supported the sentence they were tied to [3]. The answers that seemed most helpful were often the ones with the least accurate citations, a correlation of about −0.96 [3]. A 2024 audit of three answer engines found that participants flagged overly confident language in the answers, and that a large share of statements were not supported by the sources listed [8]. A page that states its claim plainly, with the source beside it, gives the engine less room to misattribute.
Map
Look at the map. For the questions your buyers ask, a written answer may sit above the links. It draws on a handful of sources, and it shows some more than others [1] [8].
Find your place. Check whether your page is among the sources, how much of the answer comes from it, and whether the sentence citing you is actually supported by your text [1] [3].
Align. Put the direct answer in one clear passage [5]. Add the citation, the quotation or the number that supports it [1]. Keep the tone plain, because tone was not where the gain was [1] [6]. Re-measure, because the engines move [1] [7].
References
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’24), 5–16. https://doi.org/10.1145/3637528.3671900 (preprint: https://arxiv.org/abs/2311.09735)
- Brin, S., & Page, L. (1998). The anatomy of a large-scale hypertextual Web search engine. Computer Networks and ISDN Systems, 30(1–7), 107–117. https://doi.org/10.1016/S0169-7552(98)00110-X
- Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. Findings of the Association for Computational Linguistics: EMNLP 2023. https://aclanthology.org/2023.findings-emnlp.467/
- Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). https://arxiv.org/abs/2005.11401
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12. https://aclanthology.org/2024.tacl-1.9/
- Wan, A., Wallace, E., & Klein, D. (2024). What evidence do language models find convincing? Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), Volume 1: Long Papers. https://aclanthology.org/2024.acl-long.403/
- Puerto, H., Gubri, M., Green, T., Oh, S. J., & Yun, S. (2025). C-SEO Bench: Does conversational SEO work? NeurIPS 2025 Datasets and Benchmarks Track. https://arxiv.org/abs/2506.11097
- Narayanan Venkit, P., Laban, P., Zhou, Y., Mao, Y., & Wu, C.-S. (2024). Search engines in an AI era: The false promise of factual and verifiable source-cited responses. arXiv:2410.22349. https://arxiv.org/abs/2410.22349
Which of your claims would an answer engine find a source for, and where do you want to stand when it does? Tell us where you want to stand. Or continue with SEO, GEO, and AEO.