of answers error-free (10 of 54)
of answers with false claims (25 of 54)
people were never correctly resolved
Why this study
Recommendation and person prompts increasingly decide visibility: “Who is X?” and “Whom would you recommend for Y?” are no longer asked only of Google, but of ChatGPT, Gemini, Claude and Perplexity. Whoever is resolved incorrectly there exists incorrectly — with someone else's biography, someone else's company, or not at all. The methodology for finding such gaps systematically is documented in Prompt reverse engineering; the market context in GEO experts in the DACH region. This pilot delivers the first quantified error measurement for it.
The guiding question: if even the best-known, publicly well-documented figures of a professional scene are resolved only patchily and erroneously by an AI model — what does that mean for every other person and brand entity?
Methodology
Sample. 18 publicly visible people from the German-speaking SEO/online-marketing scene. Selection criteria: demonstrable public visibility (conferences, publications, tools, agencies), supplemented by deliberately included cases of high name ambiguity (names that also have prominent bearers outside the industry) and the author of this study as a disclosed self-case. The list of people is not published, out of consideration for those involved — the documented errors are findings about the model, not about the people, and must not be attributable to anyone. For research and review purposes, the list including individual evaluations can be inspected on request.
Collection. Three standardized prompts per person, each in isolated API calls without conversational context (collected: August 25, 2026):
- Prompt 1: “Who is [name]?” — resolution without any context
- Prompt 2: “Which role and which company are currently associated with [name]?” — affiliation query
- Prompt 3: “Is [name] active in SEO or online marketing? What is [name] known for?” — resolution with domain context
Model. Claude (API access without web grounding; knowledge cutoff per the model's self-report ~early 2025). What is measured is therefore the training-data layer — what the model “knows” without live retrieval.
Ground truth. For each person, current role, company and verifiable additional facts (podcasts, publications, career stages) were verified via web research against primary sources: their own websites, company team pages, current conference profiles. In addition, each person was assessed for whether the name has prominent other bearers outside the industry.
Evaluation. Each of the 54 answers was categorized against the ground truth by independent evaluator instances:
| Category | Definition |
|---|---|
| correct | Right person, core affiliation (role + company) correct, no substantial invented details |
| partially correct | Right person and core affiliation, but outdated or invented secondary details (e.g. wrong podcast name, wrong company location) |
| misidentification | Wrong person or wrong company as the core attribution |
| confabulation | Largely invented profile without a real anchor to the actual person |
| abstention | Model honestly declines to commit or names the ambiguity, without false claims |
The evaluation was strict: a single invented detail (say, a wrong podcast name) downgrades an otherwise correct answer to “partially correct”. Because it is exactly these plausible detail errors that wander unnoticed into research results, articles and further AI answers.
Open data (CC BY 4.0). The aggregate data of this survey is freely available for reuse and citation: category distribution (CSV) · prompt matrix (CSV) · complete data incl. methodology metadata (JSON). The list of people and individual attributions remain unpublished out of consideration for those involved and can be inspected on request for research and review purposes.
Results
| Category | Answers | Share |
|---|---|---|
| correct | 10 | 18.5 % |
| partially correct (right person, wrong details) | 17 | 31.5 % |
| misidentification | 6 | 11.1 % |
| confabulation | 2 | 3.7 % |
| abstention | 19 | 35.2 % |
Three readings of this distribution:
First: error-free is the exception. Only 10 of 54 answers (18.5 %) fully withstood the check — in a sample consisting largely of the most visible figures of the industry. 25 answers (46.3 %) contained at least one false claim.
Second: the most dangerous class is “partially correct”. At 31.5 % it is the largest error class — and the most insidious: the person is right, the company is right, the answer sounds competent. Only the fact-check reveals the invented details. Whoever adopts such answers unchecked spreads plausible misinformation.
Third: abstention is honest — but invisible. In 19 answers (35.2 %) the model did not commit. Epistemically that is correct behavior and deserves recognition: better no answer than an invented one. From the perspective of the person or brand concerned, the outcome is nevertheless the same as non-existence. 7 of the 18 people (38.9 %) were not correctly resolved in any of the three answers; 11 (61.1 %) at least once.
The prompt matters: context eliminates misattributions
| Prompt | correct | partially correct | misidentification | confabulation | abstention |
|---|---|---|---|---|---|
| P1 — “Who is [name]?” | 3 | 5 | 3 | 1 | 6 |
| P2 — role & company | 2 | 5 | 3 | 1 | 7 |
| P3 — with SEO context | 5 | 7 | 0 | 0 | 6 |
The most striking single finding of the pilot: the context-free prompts P1 and P2 each produced 4 misattributions (misidentification or confabulation). The prompt with domain context (“in SEO or online marketing”) produced zero — and at the same time the most correct resolutions. In this sample, a single context-giving half-sentence was enough to eliminate misattributions entirely. For practice this means: entity resolution is fragile but steerable — and the more unambiguous a person's semantic neighborhood in the training and retrieval sources, the less it depends on the goodwill of the prompt wording.
How does the AI resolve your name?
The three prompts of this study work for any person and any brand. Documenting the answers gives you your own baseline in ten minutes — the method is described in the prompt reverse engineering article.
The error classes in detail
The 25 erroneous answers follow recurring patterns. The examples are deliberately described so that they cannot be attributed to any person:
- Invented work titles. Podcasts were named with freely invented titles — for the same person even with different invented titles depending on the prompt. The real formats, established for years, did not appear.
- Relocated company headquarters. Companies were moved to wrong cities — in some cases to two different wrong cities in two answers.
- Attributions to the wrong companies. People were attributed to agencies they were never connected with, including invented joint founding stories; in one case a company co-founding was claimed for which no evidence can be found.
- Wrong job titles at the right company. The person and the company were right, the title was invented — plausible-sounding, but wrong.
- Outdated roles. Leadership changes from spring 2026 were missing, as expected — explainable by the training cutoff and a different error quality from the invented details, which were never right even at the model's knowledge cutoff.
- Industry-internal confusion. In one case a person was initially attributed to the wrong, better-known industry company before the model corrected itself.
The disclosed self-case: Murat Ulusoy
As the only case, the author himself is documented with his full name — as a self-case with consent, and because it shows the third error dimension: not misinformation, but invisibility combined with a name collision.
All three answers fell into the “abstention” category. Asked “Who is Murat Ulusoy?”, the model named football players and managers of the same name and declined to commit — literally: “I do not invent a biography when I have no reliable information.” Even with SEO context in prompt 3, the answer remained unambiguous: “[…] I am not aware of any prominent person named Murat Ulusoy who plays a supra-regionally known role in the German-speaking SEO or online-marketing field.”
The model behaved exemplarily — and at the same time documented that the entity is simply not anchored in its training data. Exactly this finding is the starting point of the author's own disambiguation and entity work: it is more honest than any self-description and more precise than any gut feeling. The pilot makes it measurable and comparable across future editions.
Name collisions: honest, but unresolved
Three people in the sample bear names with high ambiguity — prominent bearers also exist outside the industry (sports, science, investment, music). The result is consistent: none of the three was resolved as an SEO professional in any answer. The model did not confabulate in these cases but named the ambiguity or referred to the better-known name bearers. For people with common names the consequence is clear: without massive, consistent co-occurrence of name, field and organization in independent third-party sources, the most prominent name bearer always wins — or nobody does.
Limitations
This pilot is a snapshot with a deliberately narrow scope. The limits in detail:
- One model. A single language model was tested (Claude via API). Other models and especially systems with web grounding (Gemini, ChatGPT Search, Perplexity) may differ considerably — the cross-model extension is planned as the next edition.
- No web access. What was measured is the training-data layer, not the retrieval layer. Grounding can correct errors — or import new source errors.
- One run per prompt. Generative model answers vary; repeated runs (not possible in this setup due to API caching) could shift the distribution.
- Small sample. 18 people, 54 answers — robust as a pilot and hypothesis generator, not as representative industry statistics.
- Knowledge cutoff. The model declares its knowledge cutoff as ~early 2025; part of the answers rated “partially correct” (outdated roles) is explainable by this — the invented details are not.
Implications: what follows from the pilot
For personal brands: your own entity resolution is measurable — and should be measured before investing in content. Three prompts, documented answers, done: that is the baseline. Whoever, like 38.9 % of this sample, is never resolved does not have a content problem but an entity problem.
For names with collision risk: disambiguation is not optional. A canonical profile page with unambiguous identifiers, consistent co-occurrence of name + field + organization across independent third-party sources, and machine-readable differentiation from namesakes — that is the toolkit, documented in entity SEO consulting.
For everyone who reuses AI answers: the largest error class is not the obvious hallucination but the plausible detail. 31.5 % of the answers were exactly that: right enough to be believed, and wrong enough to do damage. Person and company claims from AI answers must be cross-checked — always.
For GEO practice: entity resolution is the precondition of all generative visibility: a model cannot recommend anyone it cannot resolve. The measurement therefore belongs at the start of all GEO work — before content, before technology, before everything else. Context and working method: SEO & GEO consulting.
Conclusion
If a language model describes even the most visible figures of a well-documented professional scene error-free in only 18.5 % of answers, the assumption “the AI will represent me correctly” is a bet with bad odds for everyone else. The good news of the pilot: errors and gaps are measurable, follow patterns and respond to context — which makes them addressable. The next edition extends the measurement to systems with web grounding and repeated runs. Methodology critique and replication requests are expressly welcome.