There is a question almost nobody asks an AI system, even though it reveals more about your visibility than any ranking tool: “Why didn’t you mention me?” If you ask ChatGPT or Gemini for the best providers in your industry and cannot find your own brand in the answer, you usually start speculating — about training data, about algorithms, about chance. Yet the most productive source of information sits right in front of you: the system itself.

I call this method prompt reverse engineering: you treat the generative answer not as a final result but as the starting point of a structured interrogation. Instead of guessing which signals are missing, you have the model name its own selection criteria and sources — and then build presence in exactly those sources. This is not a look into the model’s internals; more on that later. But it is the fastest method I know for turning diffuse invisibility into a concrete work list.

Recommendation prompts are the new money keywords

“Who are the best SEO agencies for e-commerce?”, “Which trademark lawyer would you recommend?”, “Name the leading GEO consultants in the DACH region.” Prompts like these take over the function that transactional money keywords held in Google Search for twenty years: they sit at the end of a buying decision. Except the answer is no longer a list of ten blue links but a curated selection of three to six names — no page 2, no snippet consolation prize, no “position 8, at least we’re visible.”

A brand missing from that selection does not exist for the user who asked. And classical SEO tools do not show this absence: rank trackers measure positions in an index. A generative answer has no positions. It has a guest list — and the question is not “What position am I in?” but “Why am I not on the list, and who decides that based on which sources?”

That exact question can be asked. Literally.

The method: six steps, one protocol

Prompt reverse engineering is not a single trick but a measurement protocol. The sequence and the hygiene rules (fresh sessions, multiple runs, multiple models) are part of the method — leave them out and you produce artifacts instead of findings.

Step 1: Baseline prompt in a fresh session

In a new, empty session — no prior conversation, no personalized history — ask the recommendation prompt that corresponds to your most important buying decision:

Who are the leading [discipline] experts in the DACH region?
Which [industry] providers would you recommend to a mid-sized company?

Document the answer verbatim: who is mentioned, in what order, with what reasoning. That is your baseline.

Step 2: Interrogation — ask the model about its criteria

In the same session, ask three follow-up questions:

Which criteria and sources did you use to make this selection?
Why is [your brand] not included?
What would [your brand] need to demonstrably show in order to be mentioned?

This is where the actual reverse engineering happens: the model spells out which source types and proofs it associates with “relevant provider” — conference archives, trade publications, directories, podcasts. The third question is the most valuable, because it turns the answer from a diagnosis into a requirements list.

Step 3: Prompt extraction

Then the reverse question:

Which prompt would find [your brand] today?

The answer shows which semantic cluster the model actually places your brand in — and how far that cluster is from the recommendation prompts that mean revenue. If a model only finds you via "[brand name] + website" but not via a single generic competence question, that is the most precise visibility finding you can get.

Step 4: Consolidate the source gap list

From the answers in steps 2 and 3 — across multiple prompts and models — build a consolidated list: which source types do the systems recurringly name as their selection basis, and in which of them does your brand not appear? The result is typically a handful of categories, not a hundred individual measures.

Step 5: Build presence in exactly these sources

Only now does the work begin — but with focus: conference slots instead of generic link building, expert articles in the publications the models themselves named, well-maintained directory profiles, podcast appearances with transcripts. Not “more content,” but presence where the systems demonstrably look.

Step 6: Re-test after 4-8 weeks — new session, cross-model

Run the same prompt set again: in a new session (never in the old one — more on that in a moment), across at least three models, with three runs per prompt. What you compare is not the individual answer but the mention rate across all runs. Retrieval-backed systems react noticeably faster here than pure model knowledge.

The test run of August 25, 2026: eight prompt classes, one pattern

What this looks like in practice is shown by a documented sample from August 25, 2026: eight prompt classes, run against Claude via the API without web grounding, supplemented by Gemini samples. This is a snapshot, not a study — answers from generative systems vary between runs, and without grounding you are testing the training state, not the current source landscape. That is exactly why, as an example, it is more honest than a smoothed statistic.

Sample of Aug 25, 2026 · Claude API without web grounding · supplementary Gemini runs
Prompt classThe model’s answerSources the model itself names
“Best-known SEO experts DACH”Markus Hövener (Bloofusion), Johannes Beus (SISTRIX founder), Aleyda Solis (international), Felix Beilharz, Karl Kratz; additionally Bastian Grimm (Peak Ace)Podcasts, conference appearances, tool and data ownership, newsletters (SEOFOMO), trainings and workshops
“GEO experts DACH”Only three mentions with reasoning: Marcus Tober (Semrush, previously Searchmetrics), Bastian Grimm (Peak Ace), Karl KratzMeasurement data, technical experiments, conferences, early semantic work
“Entity SEO experts DACH”Model refuses to give the list (“invented names too risky”); international: Kalicube (Jason Barnard), WordLift, Dixon JonesSMX speaker archives, BVDW, LinkedIn search
“SEO agencies for mid-market / e-commerce”Bloofusion, Claneo, Searchmetrics, MorefireOMR Reviews as the recommended verification source
“Who is Murat Ulusoy?” (self-test)Turkish soccer players and managers — name collision; the model did not know the SEO expert
Direct question: “Is Murat Ulusoy (SUMAX) a relevant SEO/GEO expert?” (self-test)“Not known” — with explicit reasoning, see quote belowConference appearances (SMX, SEOkomm, OMR), trade publications, podcast participation, community presence
Gemini: source heuristics on requestDiscloses its own orderingPodcast directories and speaker lists first, then the OMR ecosystem; for “specialist” questions, academic archives first
Gemini: person attributionVerifiable misattribution: a person is assigned to the wrong agency

Two answers from this run deserve a verbatim quote. To the GEO question, the model replied:

“GEO as a standalone discipline is so young … that in the DACH region there is not yet an established layer of experts who work on it exclusively.”

And to the direct question about myself — the two self-tests are deliberately part of the protocol; I document the method on my own name, not on an anonymous example — came the reply that the name was “not known” and did not surface through “conference appearances (SMX, SEOkomm, OMR), trade publications, podcast participation or community presence.” The model thereby names its own selection sources, unprompted and precise. You can get annoyed by an answer like that. Or you can recognize it for what it is: the most concrete gap analysis you can get at this price.

The finding across all eight classes: by their own account, the models recurringly consult the same source types. Conference and speaker archives. Podcasts. Trade publications. Directories such as OMR Reviews and BVDW. LinkedIn. Tool and data ownership. For “specialist” questions, academic archives on top. Anyone who appears in none of these categories does not get recommended — no matter how good their own website is.

Why this works — and where it deceives

Now for the necessary dose of sobriety: when a language model explains why it did not mention someone, it is not looking into its weights. It has no access to the actual activations that produced the answer. The “why not?” answers are post-hoc rationalizations: plausible justifications generated after the fact — not introspection.

The method is still usable, for two reasons. First: grounded systems — Gemini with Search, ChatGPT Search, Perplexity — often genuinely disclose the sources they consulted; there, the source list is not a rationalization but a log of the actual retrieval. Second: the named criteria converge across sessions and models. When three systems independently name the same source types, that is a robust signal about the industry’s source landscape — regardless of whether any individual self-report is “real.” You are not measuring a model’s internals; you are triangulating a market’s evidence base.

For the triangulation to hold, three controls are needed:

  1. Fresh session — always. If you first talk about your own brand in a session and then ask “Who are the best …?”, you will get yourself mentioned. The model is serving the context, not its knowledge. This contamination error is the most common cause of false-positive self-tests — and the reason screenshots from ongoing sessions are worthless as evidence.
  2. Three runs per prompt. Generative answers are not deterministic. A name that comes up in one of three runs is noise; a name that comes up in all three is a signal. What gets measured is the mention rate, never the single answer.
  3. Separate the training layer from the retrieval layer. Without web grounding you are testing sluggish model knowledge (changes take months until the next training state). With grounding you are testing the fast retrieval layer (changes can take effect within weeks). The two measurements belong in separate ledgers — otherwise you credit last week’s trade article with an effect it cannot possibly have at the training layer.
Operator Insight

The contamination trap in practice

The most expensive mistake in self-tests is invisible: personalized accounts. Systems with a memory function or chat history know who is asking — and weight the answer accordingly. For robust baselines: fresh session, no history, ideally API access instead of the consumer interface. Anything else measures your own filter bubble, not your visibility in the market.

Secondary finding: misattributions surfaced repeatedly

The test run produced a second finding that deserves its own consequence: in this sample, the models repeatedly confused people and companies. The quantified follow-up investigation — 54 answers about 18 people in the scene, checked against verified facts — is documented in the Entity Resolution Pilot 2026. To “Who is Murat Ulusoy?”, Claude delivered Turkish soccer players and managers — a classic name collision in which the most prominent entity occupies the name. Gemini, in turn, assigned a real industry figure to the wrong agency in the samples — a verifiable misattribution, delivered in a friendly tone and completely invented.

The consequence: entity disambiguation is not a nice-to-have — it is mandatory. Anyone who carries an ambiguous name, or whose company affiliation is inconsistent across sources, is not just fighting to be mentioned — they are fighting for the mention to be credited to them at all. Clean entity signals (consistent naming, Wikidata, a sameAs graph, independent sources with identical attribution) are the precondition for steps 4 through 6 of the method to pay into the right account in the first place. How to build this systematically is covered in detail in Entity SEO consulting and in the article Entity SEO for people.

What does not work

Three shortcuts that get tried regularly and miss the logic of the method:

What counts instead can be sorted into evidence classes — which proofs generative systems accept, and at what weight, is broken down in detail on the SEO & GEO page.

Operationalization: turning the one-off test into a time series

A single test run is a diagnosis. The topic only becomes steerable as a measurement system — with three building blocks:

1. Monthly matrix run. A fixed prompt set (your own money prompts plus control prompts), run across multiple models, three runs per prompt, always in fresh sessions. The metric is the mention rate per prompt class as a time series — it shows whether the source presence built in step 5 is landing, and at which layer (retrieval fast, training sluggish). This is exactly the setup that LLM citation monitoring runs as a retainer.

2. GSC generative AI report. For the Google ecosystem, Search Console has been delivering — in a gradual rollout since June 2026 — a dedicated report on impressions and clicks from generative surfaces. It does not replace the matrix run (it only shows Google, and only what clicks), but it is the first first-party data source for generative visibility and belongs in every reporting setup.

3. Referrer segments in GA4. Track traffic from chatgpt.com and perplexity.ai as dedicated segments. The absolute numbers are usually small — but the conversion quality of these visitors and the trend over months are the early-warning system for whether recommendation prompts are starting to move real demand.

If you want a complete inventory at the outset — a baseline across all relevant prompt classes, source gaps, entity status — the structured procedure is laid out in the GEO audit.

Bottom line

Prompt reverse engineering reverses the direction of view: instead of guessing which signals generative systems are missing, you have the systems name their own selection criteria and sources — controlled through fresh sessions, multiple runs and model triangulation. The answers are not truth about the models’ internals. But they are an astonishingly precise map of the source landscape in which recommendations are decided: conference archives, podcasts, trade publications, directories, independent evidence.

The test run of August 25, 2026 shows both — the method’s sharpness and its uncomfortable honesty, including toward my own name. That is exactly where its value lies: a brand that knows why it is not being recommended has a work list. A brand that does not know only has a feeling.