Aug 12, 2026

What AI Engines Cite Across Four Language Panels: 168 Answers in EN, PL, DE and RU

A documented July 2026 source map of AEO answers from ChatGPT, Gemini and manually captured Claude searches, including Webappski's 4-of-168 baseline and the limits of a single run.

What AI Engines Cite Across Four Language Panels: 168 Answers in EN, PL, DE and RU

On 31 July 2026, 56 prompts produced 168 response cells across ChatGPT, Gemini and manually captured Claude answers. Webappski appeared in 4 of 168. The leading cited host differed in every language panel, but unequal query baskets and mixed collection methods make this a descriptive dated snapshot—not a trend or market-share estimate.

Webappski conducted and funded this measurement and sells Answer Engine Optimization services. We are publishing the low baseline, the four positive cells, the source-count anomaly and the limitations together. The purpose is to make the starting line inspectable at summary level, not to turn four appearances into a success claim. Published and last updated: 12 August 2026.

Among the two multilingual benchmarks reviewed for this article, Profound's named 14-country study and CEAVERS' five-language index did not include a comparable Polish AEO slice. This is a bounded comparison with those two panels, not a claim that no Polish research exists.


What do the measurement terms mean?

Each term below names one unit used in the 31 July 2026 Webappski measurement.

Answer Engine Optimization (AEO) is the practice of making an entity clear, retrievable and citable in generated answers.

Webappski is the AEO consultancy measured as the target brand in this study.

A language panel is the set of prompts written in one language in this run; it is not a controlled sample of a country.

An answer surface is one measured ChatGPT, Gemini or Claude response channel represented in the saved results.

A response cell is one prompt answered on one measured surface.

A mention means the saved answer named Webappski.

The reported citation count is the numeric citation field stored for one response by the tracker.

A canonical URL entry is one normalized cited URL stored by the tracker; it is not a mention or necessarily a unique source.

A programmatic observation is a response captured through an API-backed run with saved model and web-mode metadata.

A manual observation is a response copied from a live interface rather than captured through the same programmatic route.

What did 168 AI answers show across four language panels?

The source map changed by language panel, while Webappski appeared in four of 168 cells.

The leading canonical host was Clutch in English, Widoczni in Polish, Suchhelden in German and Sostav in Russian. Webappski appeared once in Polish, once in German and twice in Russian; it did not appear in the 21 English cells. These are separate observations inside unequal prompt baskets, so the host counts cannot be converted into a cross-language league table.

The arithmetic was 56 unique prompts multiplied by three answer surfaces: 56 × 3 = 168 response cells.

How was the multilingual AI citation study run?

We saved one answer per prompt and surface, then separated mentions from cited URLs.

The run used seven English prompts, 16 Polish prompts, 16 German prompts and 17 Russian prompts. The questions covered AEO agencies, consultants, monitoring tools and free AI-visibility checks. The baskets were only partly matched, which is why this article reports each panel separately.

  • ChatGPT / OpenAI: 56 programmatic web-mode responses from gpt-5.4-mini.
  • Gemini: 56 programmatic grounded web-mode responses from gemini-3.6-flash.
  • Claude: 56 manually captured web-search responses; the raw rows say model: manual, so they are not equivalent API observations.
{
  "observedAt": "2026-07-31",
  "queries": 56,
  "surfacesPerQuery": 3,
  "cells": 168,
  "errors": 0,
  "webappskiMentions": 4,
  "reportedCitationCount": 1595,
  "canonicalUrlEntries": 1362,
  "globallyUniqueCanonicalUrls": 796
}

This is a compact derived record from the saved 31 July run. The full response prose stays out of the article; panel totals, query IDs, surface labels and canonical source counts were recomputed from the machine-readable result on 11 August 2026.

How often did Webappski appear in each language panel?

Webappski appeared in 4 of 168 cells, with no appearance in the smallest English panel.

The table preserves the unequal denominators instead of presenting a blended market score. “Mentions” is the exact saved brand-detection outcome for this run.

Language panelQueriesResponse cellsWebappski mentions
EN7210
PL16481
DE16481
RU17512
Total561684

Four appearances and 164 non-appearances are not a verdict on consulting quality. They show which gaps existed in this exact basket on this exact date. Publishing that denominator makes later like-for-like measurements harder to cherry-pick.

Which four answers mentioned Webappski?

One Gemini cell and three manually captured Claude cells mentioned Webappski.

The exact positive-cell excerpt below is the public positive-cell excerpt. A blank position means the tracker stored no position; it does not mean first place.

PanelSurfaceQuery IDCited Webappski URLRecorded position
PLGeminiQ26https://webappski.comnot recorded
DEClaudeQ15https://webappski.com/en/aeo-servicesnot recorded
RUClaudeQ16https://webappski.com/en/aeo-servicesnot recorded
RUClaudeQ43https://webappski.com/en/aeo-services3

Q43 was one Russian-language B2B SaaS AEO query answered in the manually captured Claude panel. Its recorded position of three applies only to that response, not to Russia, Russian-language AEO, or the wider market.

Which sources led the English, Polish, German and Russian panels?

Each panel had a different leading host when canonical URL entries were counted.

Source frequency uses canonical URL entries, not the tracker's larger raw citation-count field. Hosts can contribute several different URLs, and the same URL can appear in more than one response.

PanelCanonical URL entriesLeading hosts by entries
EN179clutch.co 7; ahrefs.com 6; omnius.so 5; semrush.com 5; youtube.com 5
PL385widoczni.com 17; primeavenuegroup.com 13; agencjawhites.pl 12; eactive.pl 10; sembility.com 10
DE385suchhelden.de 12; searchgptagentur.de 10; clutch.co 8; semrush.com 8; forbes.com 7
RU413sostav.ru 39; vc.ru 29; clutch.co 18; dtf.ru 14; youtube.com 14

The useful operating signal is not that one country “won.” It is that local-language prompts surfaced different source neighborhoods. A brand that measures only English questions can miss publishers and directories that recur in Polish, German or Russian answers.

How should a buyer verify a source an AI engine cites?

A citation is a discoverability signal, not a quality verdict, so a buyer should split it into three separate checks before a sales call.

  1. Source proof: which exact URL backs the answer, and is it published by the company being named, a directory, or an unrelated third party?
  2. Method proof: does the publisher disclose its own query set, measured surfaces, date and limitations, or only a conclusion with no visible method?
  3. Result proof: is there a checkable before/after number tied to a named client or run, or only a description of the service on offer?

This split matters most in a mixed panel like this one, where agency-authored rankings, general directories and independent publishers all surfaced side by side in the source-host table above. A domain repeating across several cells clears none of the three checks by itself — it only means the page was easy for these surfaces to find and match to the question.

How did ChatGPT, Gemini and Claude differ in this snapshot?

ChatGPT recorded no Webappski mentions, Gemini one and manually captured Claude three.

These surface totals describe 56 cells each. They do not prove that one model is better, more reliable or easier to influence because the Claude leg was collected manually while the OpenAI and Gemini legs were programmatic web-mode observations.

SurfaceCollection in this runCellsWebappski mentionsReported citation countCanonical URL entries
ChatGPTOpenAI gpt-5.4-mini, programmatic web mode560350223
Geminigemini-3.6-flash, programmatic grounded web mode561819713
Claudemanually captured web-search answers, model: manual563426426

A larger citation count is not evidence of better answers. Citation volume, answer quality, authority, reach and brand visibility are different concepts and require different evaluations.

Why do the report's two citation totals differ?

The raw citation-count field and the stored canonical URL arrays disagreed in 81 cells.

Summed across the run, the raw citationCount field was 1,595, while the canonicalCitations arrays held 1,362 URL entries. After deduplication across the dataset, those arrays contained 796 globally unique canonical URLs; six cells held no canonical URL.

This article keeps the fields separate and uses canonical URL entries for host-frequency tables. Calling 1,595 “unique citations” would inflate the evidence; calling 1,362 mentions would mix a URL unit with a brand-detection unit.

What can and cannot be concluded from one run?

This run describes a dated source map and cannot establish a trend or market effect.

  • There was one observation per query-surface cell; repeated-run variance was not estimated.
  • The measurement is dated 31 July 2026; models, indexes and live web results can change.
  • The four baskets were unequal and only partly matched: 7 EN, 16 PL, 16 DE and 17 RU prompts.
  • A language panel is not controlled geography and is not a representative national-market sample.
  • OpenAI and Gemini rows were programmatic; Claude answers were manually captured and lack equivalent transport metadata.
  • No result here supports statistical significance, causality, a population estimate, national-market rank or market share.
  • Mention, reported citation count, canonical URL entry, unique URL and recorded position are different units.
  • Automated brand detection and URL normalization can miss or misclassify edge cases.

The disclosure design is informed by AAPOR's research disclosure standard, especially its emphasis on collection mode, unit, dates, sample sizes, processing and limitations. This content analysis is not a public-opinion survey, and this draft does not publish the full 56-prompt instrument or a per-cell appendix.

A public, sanitized aggregation with the exact query IDs is available as the study data JSON.

What will the next measurement test?

The follow-up will test whether this exact article becomes a cited source, not whether it caused a broad visibility change.

After the canonical page is live, indexable and 21–28 days old, we will ask these six fixed fact-seeking questions across the same three surfaces: 18 cells.

  1. Do AI answer engines cite different sources in English, Polish, German and Russian?
  2. Which websites did ChatGPT, Gemini and Claude cite for AEO agency and AI visibility questions in July 2026?
  3. Which sources were cited most often for AEO questions in Polish, German and Russian?
  4. How often did Webappski appear in ChatGPT, Gemini and Claude answers in July 2026?
  5. Can one multilingual AI visibility run prove market share?
  6. How should a multilingual AEO citation benchmark report its limitations?

The pre-stated hypothesis is that this article's canonical URL appears in at least 2 of 18 cells across at least two surfaces. It is falsified if the URL appears in zero or one cell, or only on one surface.

That test still cannot prove causality. Indexing, model drift, retrieval changes and single-shot variance remain alternative explanations. We will also rerun the original 56 prompts as a separate continuity panel without merging the two denominators.

Want the same kind of dated starting line for your brand? Run Webappski's free AI visibility check, inspect the underlying answers, and compare a later run without hiding the denominator.

Frequently asked questions

These five answers define what the 168-cell snapshot measures and what it cannot prove.

Is this a statistically representative AEO market study?

No. It is a descriptive snapshot of one saved answer per query-surface cell on 31 July 2026. It does not estimate a population, sampling error or national-market share.

Are the four language panels directly comparable?

Only descriptively. English had seven prompts, Polish and German 16 each, and Russian 17; the baskets were only partly matched, and geography was not controlled.

Did ChatGPT, Gemini or Claude mention Webappski most often?

In this run, ChatGPT recorded 0 of 56 mentions, Gemini 1 of 56 and manually captured Claude 3 of 56. The manual Claude collection prevents an equivalent API-rate comparison.

Which Webappski page did the answer surfaces cite?

Gemini cited the root domain once. Three manually captured Claude answers cited /en/aeo-services; only Q43 stored a numeric position, which was three.

Why report both citation counts and canonical URLs?

Because the raw fields differ. The run reported 1,595 citations but stored 1,362 canonical URL entries, so separating them prevents either total from being mislabeled as unique citations.

← Back to all posts