Sep 1, 2026

Do ChatGPT, Gemini and Claude Cite the Same Sources? 616 Hosts, 14 Shared by All Three

The follow-up our four-language snapshot promised, narrowed to one language: 267 saved Polish-language answers, 39 identical questions per engine, and the first source-overlap metric we have published — with the coverage asymmetry that partly explains it.

Do ChatGPT, Gemini and Claude Cite the Same Sources? 616 Hosts, 14 Shared by All Three

In 267 saved Polish-language answers measured by Webappski, three AI engines cited 616 distinct hosts — only 14 shared by all three. 504 hosts (81.8%) were cited by exactly one engine. Webappski ran and funded this measurement and sells the same services it counts; we appeared in 21 of the 267 observations.

This is the measurement our own four-language snapshot said it would run next. That article, published on 12 August 2026, closed with a fixed list of follow-up questions, and the first one read: “Do AI answer engines cite different sources in English, Polish, German and Russian?” We are answering a narrower and harder version of it first — inside one language, where all three engines were asked the same 39 questions, how much do their source sets actually intersect?

Two things separate this issue from that snapshot. The Polish corpus is roughly 5.5 times larger — 267 observations against the 48 Polish response cells of the July run — and the overlap is computed rather than described: the earlier article counted mentions and citation volume, but never measured how much the engines agreed on sources. Everything below is the Polish slice of a cross-section frozen on 1 September 2026. It is a dated description, not a trend, not a market-share estimate, and not a ranking of agencies.

Poland is the test bench here, not the subject. It is the slice we picked for the job: three engines, the same 39 questions put to each of them, and a market language that is not English — which is exactly the condition an English-only measurement never tests.


What does source overlap between AI engines mean?

Source overlap is the share of cited hosts that more than one engine cites when answering the same questions.

Each term below names one unit in this measurement. Mixing them is the most common way an AI-visibility number stops meaning anything.

An answer engine is a system that answers a question in prose and attaches sources. Here that means ChatGPT (OpenAI), Gemini (Google) and Claude (Anthropic).

An observation is one saved answer from one engine to one question in one run. This slice holds 267 of them.

A citation instance is one occurrence of a cited URL inside a saved answer. One host can contribute many instances; this slice holds 2,006.

A distinct host is one registrable domain, counted once no matter how many URLs or answers it appeared in. This slice holds 616.

The Jaccard index is shared hosts divided by the union of two engines' hosts. A value of 1.0 means identical source sets; 0 means the two engines had no host in common.

Coverage asymmetry is the case where two engines return very unequal numbers of sources. It pushes overlap down mechanically, without either engine disagreeing about which sources are authoritative.

A publication floor is the minimum number of observations a surface needs before we report it at all. One measured surface in this slice carried a single observation and was excluded from every aggregate, including the overlap counts. This issue is about three engines, not four.

How much did three engines agree inside one language?

They shared 14 hosts out of 616, and 504 hosts (81.8%) were cited by exactly one engine.

Pairwise, no two engines came close to the same source set. The union of any pair ran into the hundreds of hosts while the intersection stayed in the tens.

Engine pairHosts cited by bothJaccard index
Gemini ↔ Claude820.164
Gemini ↔ ChatGPT340.071
ChatGPT ↔ Claude240.065

The band is 0.065 to 0.164. For orientation: two engines drawing from an identical source pool would score 1.0, and two engines with nothing in common would score 0. What this does not establish is that any engine is more correct, that the pattern would hold on a different question set, or anything at all about change over time — this dataset has no time axis by design.

The record the numbers come from is machine-readable, and this is the part of it under discussion:

{
  "slice": "pl",
  "observations": 267,
  "distinctQuestions": 39,
  "citationInstances": 2006,
  "enginesCompared": ["gemini", "openai", "anthropic"],
  "distinctHostsAcrossComparedEngines": 616,
  "hostsCitedByEveryComparedEngine": 14,
  "hostsCitedByExactlyOneEngine": 504,
  "shareCitedByExactlyOneEngine": 0.818,
  "pairOverlap": [
    { "pair": ["gemini", "anthropic"], "sharedHosts": 82, "jaccard": 0.164 },
    { "pair": ["gemini", "openai"],    "sharedHosts": 34, "jaccard": 0.071 },
    { "pair": ["openai", "anthropic"], "sharedHosts": 24, "jaccard": 0.065 }
  ]
}

The excerpt is trimmed and rounded to the fields discussed here. The full artifact is not published — see the methodology section for why.

Which 14 hosts did all three engines cite?

Fourteen hosts survived all three engines: eight agency sites, three AI-visibility specialists, two global SEO platforms and one directory.

  • delante.pl, widoczni.com, solv.pl, silence.pl, westom.pl, rodin.pl, double-digital.pl, growpad.pro — agency and consultancy sites, the Polish ones cited on their .pl properties
  • aivisible.pl, aeovision.ai, aeoengine.ai — shops and platforms whose public positioning is AI-search visibility specifically
  • ahrefs.com, semrush.com — global SEO platforms
  • clutch.co — the only directory in the shared set

Three details in that list are worth more than the list itself. First, the most-cited host in the whole Polish slice is not in it: agencjawhites.pl took 66 citation instances across 26 questions and still failed to appear in all three engines, while widoczni.com, second at 61, did. Being cited most is not the same as being cited everywhere, and only the second of those two properties survives a change of engine.

Second, the shared Polish hosts are the local properties. Delante operates both delante.pl and an international delante.co, and the host the engines converged on in Polish answers is the .pl one. Third, a shared host is not a recommendation: clutch.co is in this set because all three engines reached for the same directory at least once, not because the directory earned agreement about anything.

How much of that gap is coverage asymmetry rather than disagreement?

A substantial part of it is asymmetry: one engine returned 350 distinct hosts in this slice and another returned 160.

This is the single most important caveat on the 81.8%, and it has to be read before the headline number is used for anything. The engines were not equally generous with sources: Gemini attached roughly fourteen citations to an answer and reached 350 distinct hosts, while ChatGPT's 308 citation instances across the whole slice landed on 160. Sets of that size cannot overlap much whatever either engine believes.

EngineObservationsQuestionsDistinct hosts citedCollection in this slice
Gemini9039350programmatic, API
Claude863923275 of 86 pasted in by hand
ChatGPT9039160programmatic, API

The concentration inside the smallest set matters as much as its size. Of ChatGPT's 308 citation instances in this slice, 51 point at one host — clutch.co — and they come from just 7 of the 39 questions. Claude's leg carries 232 hosts across 443 citation instances. So part of what the overlap metric reports as “the engines disagree” is really “one engine cites widely, one cites narrowly and repetitively”.

The Claude column has a second problem we would rather state than have a reader find. Our own published Polish methodology note says the hand-collected Claude leg normally yields zero URL citations, because the transcript we can save does not carry them; in this corpus the same leg carries 232 hosts and 443 citation instances. Citation capture on the manual leg is therefore inconsistent between collection sessions, which for engine-to-engine comparison is worse than a uniform zero would be. Any conclusion that leans on Claude's source set specifically should be treated as unsupported.

What survives the caveat is the operational point, and it survives it intact: whichever mechanism produced the divergence, a company that checked one of these engines and stopped would have seen a source landscape that the other two largely did not share.

Would an English-only check have shown you the same web?

No — in this Polish slice the largest top-level domain was .pl at 40.5% of citations, while our English slice runs 69.2% .com.

This is the one cross-language comparison this issue publishes, and it is the reason a Polish measurement is worth an English-speaking reader's time. The host mix an answer engine draws on is not a constant of the engine; it moves with the language of the question.

SliceLargest TLDSecond
Polish-language answers.pl — 40.5%.com — 36.2%
English-language answers.com — 69.2%not reported here

For an international seller the practical reading is narrow and concrete. If your category is answered in a local language somewhere you sell, the pages competing for the citation are disproportionately local-domain pages, and a visibility report built on English questions will not contain them. That is a coverage gap in the measurement, not a conclusion about the market.

Do directories carry the answer in a non-English market?

Not in this slice: directories, marketplaces and review platforms took 3.8% of citations — 77 of 2,006, across 12 of the 39 questions.

One directory accounts for most of that. clutch.co was cited 55 times in the Polish slice, and 51 of those came from ChatGPT, against 3 from Claude and 1 from Gemini, spread over 7 questions. That is one engine's habit on a handful of questions, not a market-wide pattern — and it is the same concentration that depresses the overlap metric two sections above.

The bucket that dominates instead needs stating carefully, because it is partly true by construction: 59.9% of citations in the Polish slice go to domains owned by companies the engines also name by brand. A domain lands in that bucket because the answer names its brand, so the number describes how our classifier works at least as much as it describes engine behaviour. It is not evidence that engines prefer any category of site, and we do not present it as such.

Nor is 3.8% a verdict on directory listings. Share of citations is not share of value, and the same directory sits third by absolute count in this slice. Which source types get cited in Polish answers was the subject of an earlier issue — Jakie źródła AI cytuje po polsku (in Polish) — and this issue does not restate its conclusions.

Which brands did the engines name most often, and why is that not a ranking?

Delante and Agencja Whites tie at 77 mentions each, and the list mixes agencies with software vendors.

Mention frequency here means one thing only: how often a name appeared in the answers to our own 39 Polish questions. It is not market share, revenue, client outcomes or quality.

BrandMentionsQuestions it appeared inWhat it is
Delante7726SEO/SEM agency, Kraków
Agencja Whites7725digital agency, Warsaw (Whites Sp. z o.o.)
Widoczni6323SEO agency
EACTIVE4221SEO agency
Otterly.AI309AI-search monitoring tool
Sembility2915360 marketing / SEO agency
Webappski (us)2121 of the 267 observations named us

The top of that table is an exact tie, and the tie is the interesting part. Delante and Agencja Whites both land on 77; the order they print in is an alphabetical tiebreak on the canonical entity key, not a rank, and the tool now marks tied entities explicitly so the ordering cannot be read as a result. Delante's 77 came from 26 of the 39 questions and Whites' from 25 — a real difference in spread, and still not a ranking of either firm. Whites' figure is itself two spellings folded together, “Agencja Whites” (42) and “Whites” (35).

The presence of Otterly.AI in that list is the cleanest argument against reading it as an agency league table. Otterly.AI is a monitoring product, not an agency — it sells a subscription that watches AI answers. An engine naming it in response to a buyer question is answering a different question than “which agency should I hire”, and any list built from raw mention counts will keep mixing the two. A Polish-language look at the agency question specifically is an earlier issue — Agencje AEO w Polsce (in Polish).

What had to be fixed before these numbers were safe to publish?

Three defects in our own pipeline, all caught before publication — and one of them changed the answer at the top of the table.

Duplicate inflation. The discovery glob was matching byte-identical copies of the same saved answers living in a backup tree, a git worktree and a same-day baseline folder. In the Polish slice that meant 369 observations before deduplication and 267 after; corpus-wide, 417 duplicated observations were dropped. This was not cosmetic: on the inflated corpus the Polish top ten opened with Agencja Whites ahead of Delante. After deduplication the two are tied, which is what this article reports.

Alias folding. Engines write the same company several ways — “Agencja Whites” and “Whites”, “EACTIVE” and “Eactive”, “Otterly.AI” and “Otterly”. A transparent fold (case, diacritics, punctuation, legal suffixes, a trailing “AI”) with the merges printed alongside the counts collapses spelling variants into canonical entities. Near-collisions were deliberately not merged and stay in the tool output as a separate list, because a vendor and its product are different entities: Semrush is not Semrush AI Visibility Toolkit.

The fold that broke on non-Latin scripts. The first version of that fold stripped every non-ASCII letter before comparing names. On a Cyrillic slice this reduced 15 unrelated Russian agency names to the same empty key and printed them as a single phantom market leader with 49 mentions. It was caught by running the tool on a slice it was never designed for, fixed with a Unicode letter-and-number fold, and locked with a regression test. If a vendor hands you a leaderboard for a market that does not write in Latin script, the question to ask is how they fold names — this class of bug is silent and it always produces a confident-looking winner.

How was this measured, and what is not in it?

Nine measurement dates between 19 April and 31 August 2026, 39 Polish-language questions, three engines, one saved answer per cell.

  • Slice rule: an observation is Polish if its question language is Polish or its target market is Poland — 154 by language only, 27 by market only, 86 by both.
  • Not one instrument. Across the whole corpus these runs used 14 distinct provider-and-model combinations, and the question grid was rewritten more than once. That is why there is no “up” or “down” anywhere in this article: a delta would measure our own tooling, not the engines.
  • Collection modes differ. The Gemini and ChatGPT legs are programmatic API runs. 75 of the 86 Claude observations were pasted in by hand from a live interface, which is a different instrument wearing the same column heading.
  • One surface was excluded. A fourth measured surface carried a single observation in this slice, below our publication floor, and appears in no number here.
  • One observation per cell. Repeat-run variance was not estimated, so nothing here supports a significance claim, a population estimate or a market share.
  • Our own numbers are in the corpus. We appear 21 times in 267 observations, counted by the same rules as everyone else and not adjusted.

Reproducibility, in place of credentials: the aggregates come from a single script that reads the saved answer files and recomputes every figure with no network calls and no model calls. It is written to ship with our open-source tracker, aeo-platform, but it is not in the published npm build as of 1 September 2026, so an outside reader cannot run it today. Saying so is more useful than a claim of independence nobody can check.

Two things are deliberately not published. The question instrument itself stays internal, as it did in the July snapshot. And the raw dataset stays internal for a harder reason: its provenance blocks name client domains and local file paths, so the file cannot be released as-is. The aggregates in this article contain no client data — 11 client runs and 225 observations were excluded from the corpus before any figure was computed, by an allow-list of our own domains rather than a deny-list of client names.

For disclosure we work against AAPOR's disclosure standard, which asks a researcher to name the sponsor and the party who conducted the work, the collection mode and dates, the sample sizes, the processing, and the design limitations. This article meets those on sponsor, mode, dates, sizes, processing and limitations. It does not meet the standard on exact question wording, and a content analysis of machine answers is not a survey of a population in the first place. We would rather state the gap than imply a rigour we did not buy.

What should a company outside Poland do with a Polish measurement?

Treat it as a test of your measurement, not as news about Poland.

The finding that travels is methodological: a visibility result produced on one engine, in one language, is a result about that engine and that language. Five specific consequences follow, and none of them requires caring about the Polish market.

  1. Do not blend engines into one score. The two programmatic engines here, ChatGPT and Gemini, shared 34 cited hosts — a Jaccard overlap of 0.071. A single blended number hides which engine produced it.
  2. Check the denominator before you call a gap a problem. An engine that names you less may simply cite fewer sources for everyone: 350 hosts versus 160 in the same 39 questions.
  3. Ask a vendor which engines a claim covers, and how each leg was collected. An API run and a hand-pasted transcript are different instruments even when they sit in adjacent columns.
  4. If you sell into a non-English market, measure in that language. The local-domain share was 40.5% there against 69.2% .com in our English slice; English questions will not surface the pages you are competing with.
  5. Ask how a leaderboard folds names. Alias handling decided the top of our own table twice — once through duplicates, once through a fold that erased non-Latin characters.

If you want the same kind of dated starting line for your own brand, our free AI-visibility check runs your category questions through ChatGPT and Gemini in about 30 seconds with no email and no signup, and shows which rivals were named instead of you; the full three-engine audit, Claude included, is available on request. The point of this issue is that two engines are already more informative than one — and that whoever measures you should tell you which ones they asked.

Frequently asked questions

Six answers on what the 267-observation Polish slice measures and what it cannot support.

Does this prove ChatGPT, Gemini and Claude disagree about which sources are authoritative?

No. Part of the divergence is coverage asymmetry — 350 distinct hosts from one engine against 160 from another, with 51 of that smaller engine's 308 citation instances landing on a single directory. The safe statement is that their source sets barely intersected, not that they judged authority differently.

Is checking one AI engine enough to measure AI visibility?

Not on this evidence. Two engines answering the same 39 questions shared only 34 cited hosts, a Jaccard overlap of 0.071, and 81.8% of the 616 hosts were cited by exactly one engine. A single-engine check describes that engine.

Why does this issue cover three engines and not four?

A fourth measured surface carried one observation in the Polish slice, which is below the floor we publish at. Rather than report a number built on a single answer, we excluded it from every aggregate and say so.

Is this a ranking of Polish agencies?

No. It counts how often a brand was named in answers to our own 39 questions; the top two are exactly tied at 77 mentions, the list includes software vendors alongside agencies, and frequency says nothing about the quality or the results of any firm.

Can I download the raw dataset?

No. Its provenance blocks carry client domains and local file paths, so the file is not publishable as-is. The aggregates quoted here exclude client runs entirely — 11 runs and 225 observations were removed before any figure was computed.

Does a higher citation count mean an engine gives better answers?

No. Citation volume, answer quality, authority and brand visibility are different concepts. In this slice the engine that reached the most distinct hosts is simply the one that attaches the most links to an answer, which is a behaviour, not a verdict.

← Back to all posts