If your brand is not cited inside ChatGPT, it is missing from a growing share of informational search. According to OpenAI, ChatGPT had more than 900 million weekly users in February 2026. As early as July 2025, Sam Altman cited more than 2.5 billion prompts per day (TechCrunch). OpenAI does not publish how many of them trigger a web search. Many brands still treat ChatGPT as a fringe channel, and SEO teams rarely have a dedicated playbook for it.
This guide closes that gap. It describes how ChatGPT technically constructs answers, which signals drive citation probability, and the eight measures that have measurably lifted citation rates within three to six months in advisory practice.
How does ChatGPT choose the sources for its answers?
ChatGPT picks sources for web answers through its own search infrastructure and search partners. According to OpenAI, ChatGPT Search has used third-party search providers and partner content since its launch on October 31, 2024; since December 16, 2024, web search has been open to all logged-in users. The OpenAI Help Center names Bing as the search engine that receives de-identified queries. On top of that, the OAI-SearchBot crawler collects websites for the search feature; pages that block it in robots.txt are not shown in search answers, per the OpenAI crawler documentation. OpenAI does not disclose Bing's share of cited URLs or how re-ranking is weighted. With more than 900 million weekly users, this is no longer a fringe channel. Three duties follow that are documented: secure Bing indexation, allow OAI-SearchBot, and write key statements so a single paragraph works as evidence without context. Everything else, such as how individual passages are weighted, is test observation, not an OpenAI statement.
Our working model of the ChatGPT retrieval pipeline (reconstructed from behavioural tests, not confirmed by OpenAI, evidence level D) has four steps. (1) Query interpretation by the language model, generating one or more sub-queries. (2) Retrieval through search partners such as Bing, supplemented by OpenAI's own search crawler OAI-SearchBot. (3) Re-ranking of the top-N results for semantic relevance to the full query. (4) Answer synthesis with explicit or implicit citation of the highest-scoring chunks.
For SEO, this is a structural shift: document ranking no longer decides — chunk quality does. A page in position one is not necessarily cited; a structurally clean passage from position eight may well be. The optimization layer moves from the page to the paragraph.
dominates ChatGPT retrieval: in our citation samples most sources come from Bing-indexed URLs (own observation, evidence level C; OpenAI publishes no shares)
tokens per chunk: our working hypothesis for highly citable passages (not a vendor figure)
of signals in our working model of source selection (own framework)
The six visibility layers for ChatGPT
Layer 1 — Bing indexation. Mandatory, not optional. No retrieval without it. Set up Bing Webmaster Tools, submit an XML sitemap, activate the IndexNow protocol. Many brands that are excellently indexed in Google have substantial gaps in Bing: in our audits, 15 to 40 percent of URLs were often missing (own observation, evidence level C, see benchmarks).
Layer 2 — OpenAI crawler access. Declare it explicitly in robots.txt. According to OpenAI, three bots matter: GPTBot (training foundation models), OAI-SearchBot (inclusion in ChatGPT search) and ChatGPT-User (fetching during user actions, where robots.txt rules may not apply). The default recommendation for B2B brands: allow all of them. For premium content publishers: block GPTBot selectively, allow OAI-SearchBot.
Layer 3 — Passage-level citability. Chunks of 200-400 tokens with a claim-evidence structure in the first sentence. No anaphoric pronouns pointing back to earlier paragraphs. Explicit entity mentions instead of implicit references. Concrete numbers with sources instead of vague phrasing. In our benchmarks, these chunk properties correlate most strongly with actual citations (own data, evidence level C). Research supports the direction: in the GEO paper by Aggarwal et al. (KDD 2024), adding citations, quotations and statistics raised visibility in generative answers by a relative 30 to 40 percent.
Layer 4 — Entity resolution. ChatGPT must be able to interpret the brand as an unambiguous entity — not as a name-collision candidate. That requires a Wikidata item with referenced properties, a Schema.org graph with @id coherence, a sameAs cluster across authoritative third-party profiles, and consistent attributes (role, industry, location, founding year). Without that resolution, the brand is either not cited at all or confused with a foreign entity.
Layer 5 — Schema.org JSON-LD. For ChatGPT, structured data is not a rich-result signal but a semantic substrate. Article schema with author-@id pointing to a Person schema, Product schema with brand-@id pointing to Organization, FAQPage with explicit claim-answer structure. Schema implementation provides the foundation.
Layer 6 — llms.txt and brand cohesion. Complementary: an llms.txt file with a structured Markdown summary of the most important pages. Not officially confirmed as a signal by any major provider; Google Search ignores llms.txt according to the Google guide to optimizing for generative AI (as of July 10, 2026). Implementation cost is low; a measurable effect for ChatGPT has not been shown so far. Plus consistent brand maintenance on LinkedIn, Crunchbase, GitHub and YouTube, all sources that frequently show up as context in our tests.
| Bot | Function | Recommendation | Consequence if blocked |
|---|---|---|---|
| GPTBot | OpenAI training crawler | Allow (publishers: weigh it up) | Content not used to train future models |
| ChatGPT-User | Fetching during user actions in ChatGPT and custom GPTs | Allow | robots.txt may not apply, per OpenAI; no effect on search inclusion |
| OAI-SearchBot | Crawler for ChatGPT search | Allow | Page not shown in ChatGPT search answers |
| Bingbot | Microsoft index (search partner, per OpenAI) | Mandatory allow | Likely much less web retrieval in ChatGPT (hypothesis, evidence level D) |
Is your bot configuration correct?
A 30-minute live check of your robots.txt and Bing indexation against the four critical crawlers — clear go/no-go list instead of guesswork.
Passage engineering: a concrete example
Weakly citable text (typical for B2B content): "Our platform offers a comprehensive solution for customer service. It combines various features that help teams support their customers better. Many companies use it to optimize their processes."
Rewritten for strong citability (fictional pattern; replace the bracketed placeholders with verified values): "[Brand] is a customer service platform that integrates ticket management, chat, AI agents and analytics in a single interface. According to [annual report, year], [number] companies in [number] countries use the platform, including [reference customer A] and [reference customer B]. [Analyst firm, study, date] rates it as [rating] in the [category] segment."
The difference is structural: the second text is self-contained (it works without context from other paragraphs), names the entity explicitly, delivers concrete numbers with sources and carries a claim-evidence pairing in its first two sentences. In our citation tests these patterns are cited most often (own observation). What matters is that every number in the final text comes from a verifiable primary source.
The 90-day protocol for ChatGPT visibility
Month 1 — audit and foundation. Bing-indexation check across all relevant URLs, GPTBot-access audit in robots.txt, entity-maturity check (Wikidata status, schema-graph coherence) and a baseline measurement of current ChatGPT citation across 1,500 prompts (brand, category, long tail).
Month 2 — passage rewrites and schema. Walk through the top 30 URLs with the highest citation potential, refactor chunks to 200-400 tokens, introduce a claim-evidence structure, make them self-contained. In parallel, roll out a Schema.org graph with @id coherence, validate the markup and optionally deploy llms.txt.
Month 3 — corroboration and entity reinforcement. Maintain the Wikidata item (references from independent secondary sources), secure three to five trade-media features with consistent fact fingerprints, set up author entities for writers. In parallel, run a second citation-rate measurement. In our projects, the uplift in weeks ten to twelve was mostly 30 to 80 percent over baseline (own data, evidence level C, no promise of results; methodology in the LLM Citation Benchmark).
What classical SEO no longer delivers
Three misconceptions that persist stubbornly in practice. First: "We rank position one in Google, so ChatGPT will cite us." Risky: OpenAI names Bing as the search partner behind ChatGPT web search, not Google. A Google ranking does not carry over automatically. Second: "We have strong backlinks, that is enough for LLMs too." Backlinks are entity signals, but without structured data and passage quality they remain unused. Third: "A solid FAQ is enough." FAQs are one signal, but without schema and without passage engineering on the main content they stand alone.
ChatGPT SEO is a discipline in its own right within the GEO stack. Classical SEO remains the foundation — but it is no longer sufficient. Ignore the channel and you lose informational demand. Address it systematically and you build a channel where, in our assessment, competitive density is still lower than in the classical Google SERP.
What measurement looks like: KPIs for ChatGPT SEO
Four primary metrics from LLM citation monitoring practice. (1) Citation rate: share of prompts in the tracking matrix in which the brand is cited. (2) Position in answer: position of the citation inside the ChatGPT response (top, middle, end). (3) Source-origin breakdown: which URL was cited — own domain, third party, Wikipedia. (4) Competitor share of voice: share of citations against competitors for the same prompts.
Secondary KPIs: entity-resolution rate (is the brand correctly recognised as a unique entity?), hallucination rate (are false attributes assigned?), source freshness (how current are the cited sources?). These secondary metrics explain the primary movements — when citation rate falls, the cause is usually entity-resolution breakage or content aging.
Bottom line: ChatGPT SEO is not an add-on
Treating ChatGPT visibility as a side task of the SEO team systematically underestimates the channel. The retrieval mechanics are different, the KPIs are different, the training and inference cycle is different. Done right, you build a channel that classical SEO does not cover, and where competitive density is, in our assessment, still lower than in the classical Google SERP.
The structural lever is not more content but better-structured content — plus a clean entity architecture plus Bing-indexation hygiene. The 90-day protocol above provides the sequence. Continuous citation monitoring makes the effect visible.
Sources
- OpenAI: Scaling AI for everyone (February 2026)
- TechCrunch: ChatGPT users send 2.5 billion prompts a day (July 21, 2025)
- OpenAI: Introducing ChatGPT search (October 31, 2024)
- OpenAI Help Center: ChatGPT search for Enterprise and Edu (accessed September 23, 2026)
- OpenAI: Overview of OpenAI Crawlers (accessed September 23, 2026)
- IndexNow: IndexNow protocol (accessed September 23, 2026)
- Aggarwal et al.: GEO: Generative Engine Optimization (KDD 2024, arXiv November 2023)
- Google Search Central: Optimizing for generative AI features (July 10, 2026)