Since December 16, 2024, ChatGPT search has been available to every logged-in user, and since February 5, 2025 to everyone, no sign-up required. According to OpenAI, more than 900 million people used ChatGPT every week as of February 2026. For brands, that means a search and recommendation channel of considerable size that works by different rules than Google's classic results list.
The decisive question for every 2026 brand strategy: how does my brand get selected as the preferred answer by ChatGPT (and Claude, Perplexity, Gemini)? This article outlines the technical playbook beyond the usual buzzword layer.
How do ChatGPT, Gemini and Perplexity choose their sources?
ChatGPT, Gemini, Claude and Perplexity base their answers on two things: knowledge learned in training and web pages they retrieve at the moment of the query. Whether a brand gets mentioned is therefore decided twice: over the long term in the training corpus, and in the short term during retrieval, in the search results the system assembles its answer from.
For AI Overviews and AI Mode, Google describes a “query fan-out” technique: the system issues several related searches across subtopics and combines the results into one answer. Since October 2024, ChatGPT search has shown the sources it used as links below the answer. If you are missing from those retrieved results, you cannot be cited at the retrieval layer. If you rarely appear in qualified trade sources, you are weakly anchored in model knowledge. Prompt-level SEO therefore works both layers: classic rankings and citable passages for retrieval, consistent mentions in trade sources for training.
The two mechanisms in detail:
Training-based answers
The model uses only knowledge learned during training. For queries like "Who is considered the leading SEO expert in Germany?" it draws on information frequently and consistently associated with certain names in the corpus. The decisive factor: frequency of mention in qualified sources during training.
RAG-based answers
For current queries or questions past the training cutoff date, the system triggers a web search. Retrieved results are passed as context to the model, which synthesizes the answer. The decisive factor: ranking in the retrieved sources plus the quality of extracted passages.
A complete optimization must address both layers.
The six levers of prompt-level SEO
Lever 1: passage engineering
LLMs do not extract entire pages for answers but individual paragraphs or sentence pairs. These must be structured as standalone units of knowledge:
- Clear definition in the first sentence: "GEO is the optimization of content for citation inside generative AI answers."
- One concept per paragraph: no multi-topic paragraphs.
- Quotable length: 2-4 sentences is ideal. Longer paragraphs get fragmented, shorter ones lose context.
- Self-contained: the paragraph must remain understandable without its surrounding structure.
For context: in its guide to optimizing for generative AI features (last updated July 2026), Google states that content does not need to be broken into tiny pieces or written in a special way for AI to appear in Google Search. Passage engineering is therefore not a Google requirement but our heuristic for citability across all systems: clearly bounded paragraphs are easier to reuse correctly.
Lever 2: entity-embedding signals
LLMs store entities as high-dimensional vectors. In our working model, the proximity of two entities in vector space correlates strongly with how likely they are to appear together in answers; it can only be observed indirectly through proxy embeddings. For a brand this means:
- The brand must be consistently placed close to relevant domain terms in high-quality sources.
- That proximity is built through co-occurrence in qualified text corpora (not through keyword stuffing).
- Source quality weights the strength of the embedding signal.
In practice, following this working model: a trade article in a highly authoritative outlet that mentions a person three times in close proximity to "AI SEO strategy" and "international scale" is likely to affect embedding positioning more than 50 LinkedIn posts with the same terms.
Lever 3: citation-target strategy
LLMs are trained systematically more often on certain sources than on others. OpenAI disclosed this in detail for GPT-3: according to the GPT-3 paper (Brown et al., 2020), 60 % of training examples came from filtered Common Crawl, yet Wikipedia was seen 3.4 times while Common Crawl was seen less than once. For current models, providers disclose their data mix only in broad terms. Documented or plausible sources include:
- Common Crawl: an open web archive, quality-filtered before training in the case of GPT-3.
- Wikipedia + Wikidata: structured, factual sources with high weight.
- Reddit: Google (February 2024) and OpenAI (May 2024) agreed on access to Reddit's Data API; Google explicitly mentions training.
- Own web crawls: according to the Claude Help Center, Anthropic collects training data with its ClaudeBot crawler, which respects robots.txt.
- GitHub, Stack Overflow: technical authority.
- Curated datasets: books, scientific publications, high-quality news.
A citation-target strategy prioritizes systematic presence in these sources across 12-24 months. In our experience, one-off PR pushes rarely reach training data because they are too thinly distributed.
Lever 4: quotable claims
Models prefer content in answers that paraphrases well:
- Specific numbers and percentages: "+23 % brand-search lift through…" beats "significantly increased brand awareness".
- Clear causal chains: "If X, then Y, because Z".
- Frameworks with named components: "The 5-step GEO framework consists of…".
- Named concepts: proprietary terminology unambiguously tied to your brand.
There is research behind this: in the GEO paper by Aggarwal et al. (KDD 2024), the methods "Cite Sources", "Quotation Addition" and "Statistics Addition" improved the visibility of sources in generative answers by a relative 30 to 40 %. Classic keyword stuffing brought little to no improvement.
Lever 5: a schema layer for RAG optimization
Google makes clear that generative AI features require no special schema.org markup. Structured data is still useful: it makes author, organization and terms unambiguous for machines and, in our working model, helps with entity resolution. Useful types:
Articlewith completeauthor/publisherdataFAQPagefor clearly answered questions (still a valid type, but Google has shown no FAQ rich results since May 7, 2026)DefinedTermfor concept definitionsHowTofor procedural contentOrganizationwithsameAslinking to existing profiles (Wikipedia, Wikidata, LinkedIn)
Lever 6: authority reinforcement
For the "final yards" of selection: signals that confirm to the model why this source should be preferred over alternatives. Author bylines with verifiable expertise, references to primary sources, quotes from other authoritative publications and academic references where possible.
"Prompt-level SEO is not 'writing content that sounds nice'. It is constructing content that models prefer to paraphrase — and that demands structural discipline in every single passage."
The measurement layer: running prompt audits systematically
Without measurement every optimization is speculation. How to reconstruct the underlying prompts from LLM answers is covered in Prompt reverse engineering. A solid prompt audit setup covers:
Building the prompt catalogue
- Category prompts: "Best SEO consultants for international brands".
- Problem prompts: "How do I optimize for ChatGPT visibility?".
- Comparison prompts: "SUMAX vs. [competitor]: which agency is better?".
- Brand prompts: "What does Murat Ulusoy do?".
- Long-tail prompts: specific sub-questions across the customer journey.
Cross-model testing
Every prompt is tested against at least four models: the current flagship models from OpenAI (ChatGPT), Anthropic (Claude), Perplexity (Sonar) and Google (Gemini). Results vary considerably because training data and retrieval priorities differ.
Metric setup
- Mention: is the brand mentioned at all? (Y/N)
- Position: first recommendation, middle or last?
- Sentiment: positive, neutral or critical?
- Competitive set: which other brands are mentioned?
- Citation link: is there a source link to your own domain?
typical build time for significant LLM visibility in an established category (based on our projects)
prompt universe for valid share-of-model tracking
conversion advantage of LLM-referred traffic vs. organic in our client portfolios (first-party data)
The overlooked lever: Reddit intelligence
Google and OpenAI both signed agreements in 2024 for access to Reddit's Data API, and Google explicitly mentions training. That makes Reddit a directly licensed source for several major providers. A targeted, serious presence in relevant subreddits — through expert posts, answers to concrete questions and authentic threads — can disproportionately affect LLM visibility. Important: no spam tactics. Reddit communities detect artificial activity, and negative signals carry into LLMs just as positive ones do.
Typical execution mistakes
- "AI-friendly" content spam: pages mass-produced in the hope that models will cite them. Without entity integration they remain ineffective.
- Llms.txt as strategy: according to Google's guide, Google Search ignores such AI text files; we know of no documented effect with other providers either.
- Keyword stuffing with AI terms: damages content coherence and reduces citability.
- Isolation from classical SEO: prompt-level SEO does not work without a solid classical SEO base (the RAG layer needs rankings).
- Expecting results in four weeks: training cycles take months. In our experience: 6-18 months until visible impact on model knowledge, sometimes faster at the retrieval layer.
The passage-engineering method: how paragraphs become citable
Every passage you want an LLM to cite must satisfy five properties at once. We use the QUEST heuristic: Quotable, Unambiguous, Entity-rich, Standalone, Timestamped.
Quotable — between 35 and 110 words. Shorter passages lose context, longer ones get clipped. Unambiguous — exactly one core argument per paragraph; no "on the one hand / on the other". Entity-rich — at least one named entity in the first sentence. Standalone — understandable without the preceding paragraph. Timestamped — contains a reference that signals currency (year, study, version).
A QUEST-compliant passage:
"According to the zero-click study by SparkToro and Datos (July 2024), 58.5 % of Google searches in the US ended without a click, 59.7 % in the EU. This trend, known as zero-click search, turns the search engine from a traffic channel into a perception channel."
Four entities (SparkToro, Datos, Google, zero-click search), two figures, one date, one self-contained core sentence. Around 40 words. In our internal passage-extraction tests, passages like this were cited noticeably more often than comparable narrative paragraphs with identical content but no structure (first-party, unpublished observation).
Cross-model testing: the reproducible audit
Every prompt is tested against four models and five repetitions. Under stochastic weighting: a majority mention in ≥ 3/5 runs per model is a stable signal. Less is noise.
Audit flow per prompt:
1. Prompt to the current flagship models from OpenAI, Anthropic,
Google and Perplexity
2. 5 runs each with temperature=0.7 (fixed, identical across models)
3. Response parse: brand mention (y/n), position (first mention: word index),
sentiment (small LLM as classifier), competitor-mentions list
4. Aggregation: per-model hit rate, cross-model consistency
5. Categorization:
- Stable Winner: ≥ 4/4 models with ≥ 60 % hit rate
- Asymmetric: 1-2 models dominant, others blank → corpus gap
- Invisible: < 20 % across all models → strategic gap
The asymmetric category is the most diagnostically valuable. In our working model, a brand strong on Gemini but invisible on Claude suggests the content lives in Google-indexed sources but not in the web sources other providers capture with their own crawlers (for Anthropic, ClaudeBot for training data and Claude-SearchBot for search). The asymmetry shows which distribution layer is missing.
Citation target playbook: which assets actually work
Not every piece of content can be cited. Our day-to-day work has identified five asset classes that consistently produce high citation rates:
- Original studies. At least 300 data points, clear methodology, reproducible design. In our experience, a first-party study produces more citation impact over 8-12 months than dozens of generic blog posts.
- Definitional articles. Deep pieces on a single term being debated in the industry. The goal: become the canonical definition.
- Comparison benchmarks. Options (tools, methods, vendors) compared along transparent criteria. LLMs use these as answer scaffolds.
- Decision frameworks. Named models ("SUMAX QUEST heuristic") established as industry vocabulary. Naming is citation-critical.
- Expert interviews. Primary quotes with author schema. LLMs frequently quote people verbatim — especially when source and person are clearly attributed.
relative visibility gain from citations, quotations and statistics (GEO paper, KDD 2024)
months until original studies reach full effect (based on our projects)
minimum test matrix: 5 runs × 4 models for stable signals
Schema layer as a citation lever
Schema.org is not a ranking factor, and Google says explicitly that generative AI features require no special markup. In our portfolios we still observe that pages with fully valid Article + author.sameAs + DefinedTerm are cited more often than semantically similar pages without that schema depth. That is a correlation from first-party data, not a controlled measurement.
The decisive point, in our experience: not presence but completeness and connectivity. An Article schema without an author provides less context than an Article with a fully resolved Person author including sameAs links to existing profiles such as Wikipedia or LinkedIn. The author becomes a standalone entity recognized across multiple domains.
Minimal schema for LLM optimization:
{
"@context": "https://schema.org",
"@type": "Article",
"@id": "https://example.com/article#article",
"headline": "...",
"datePublished": "2026-03-15",
"dateModified": "2026-03-15",
"author": {
"@type": "Person",
"@id": "https://example.com/#person",
"name": "Jane Example",
"sameAs": [
"https://www.wikidata.org/wiki/Q123456789",
"https://www.linkedin.com/in/jane-example"
],
"knowsAbout": ["SEO", "GEO", "Reputation Engineering", "LLM"]
},
"publisher": {
"@type": "Organization",
"@id": "https://example.com/#org",
"name": "Example Inc.",
"sameAs": ["https://www.linkedin.com/company/example-inc"]
},
"about": [
{"@type": "Thing", "name": "Generative Engine Optimization"},
{"@type": "Thing", "name": "LLM SEO"}
]
}
Tutorial: a 60-day implementation of a prompt-level SEO programme
Days 1-10 — diagnostics
Prompt audit with 150 prompts against 4 models. Classification into Stable Winner / Asymmetric / Invisible. Competitor overlap analysis: which brands dominate the Invisible prompts?
Days 11-25 — content retrofit
Convert your top 20 business-critical pages to QUEST. Each page gets at least three QUEST-compliant passages, full schema and author binding via sameAs. No new content yet — optimize existing first.
Days 26-40 — entity infrastructure
Maintain the Wikidata item, or create one only if the notability criteria are met. Prepare a Wikipedia draft (or commission externally). Optimize the LinkedIn company page. Close the sameAs graph: every profile links to every other profile.
Days 41-55 — corpus distribution
Negotiate and produce three Tier-1 guest contributions. Launch one original data study (survey, analysis or benchmark). Book one podcast appearance with a trade podcast.
Days 56-60 — re-audit + dashboard
Second prompt audit with an identical set. Document the delta per category. Build a weekly monitoring dashboard for Share-of-Model, PVI and Asymmetry Index.
How big is ChatGPT as a search and recommendation channel?
Big enough to plan for as a channel of its own: according to OpenAI, more than 900 million people used ChatGPT every week as of February 2026. The study “How People Use ChatGPT” (Chatterji et al., NBER, September 2025) already counted around 18 billion messages per week from 700 million weekly users in July 2025, roughly 2.5 billion per day.
Especially relevant for brands: according to the same study, the share of messages in which users look for information ("Seeking Information") grew from 14 % to 24 % between July 2024 and July 2025. An internal, unpublished analysis of 4,200 anonymized ChatGPT prompts found 31 % with direct commercial signals ("which is the best…", "which company…", "where can I…"). That first-party figure is not comparable with the categories of the OpenAI study, but it points the same way: recommendation questions often come right before a purchase decision.
Conclusion
Prompt-level SEO is the logical evolution of SEO in a world where search is increasingly processed generatively. It is not easy, not quick, and demands different skills than classical SEO. But organizations that approach it systematically now stand a much better chance of appearing, two years from now, inside the LLM answers that lead their customers to daily purchase decisions.
The others will develop a very concrete feeling for what "invisible in AI" means.
Sources
- Introducing ChatGPT search, OpenAI, October 31, 2024, updates of December 16, 2024 and February 5, 2025
- Scaling AI for everyone, OpenAI, February 27, 2026
- How People Use ChatGPT (Chatterji et al.), NBER Working Paper 34255, September 2025
- AI features and your website, Google Search Central, last updated December 10, 2025
- Google's Guide to Optimizing for Generative AI Features on Google Search, Google Search Central, last updated July 10, 2026
- FAQ rich result deprecation (changelog), Google Search Central, May 2026
- Language Models are Few-Shot Learners (Brown et al., GPT-3), arXiv 2005.14165, 2020
- GEO: Generative Engine Optimization (Aggarwal et al.), arXiv 2311.09735, KDD 2024
- Google expands partnership with Reddit, Google, February 22, 2024
- OpenAI and Reddit Partnership, OpenAI, May 16, 2024
- Does Anthropic crawl data from the web, and how can site owners block the crawler?, Claude Help Center, Anthropic
- 2024 Zero-Click Search Study, SparkToro/Datos, July 2, 2024
- Wikidata:Notability, Wikidata, help page