Study · Kyiv · Warsaw · Vilnius · August 2026
Does the language you ask in change which businesses AI recommends?
We asked five AI answer engines the same twenty dental questions about three capital cities — once in the local language, once in English — and repeated the whole thing three times in each. 1,800 answers. The recommended clinics differ by language in all three markets, beyond what the engines' own instability explains. The effect is concentrated below the top of the leaderboard, and how big it looks depends on how you count.
Method, before the numbers
- Markets
- Kyiv (Ukrainian / English), Warsaw (Polish / English), Vilnius (Lithuanian / English). Three countries, so the result reads as a property of answer engines rather than a quirk of one language.
- Questions
- 20 buying questions a real patient would type, written natively in each language and translated 1:1. The same twenty questions in every market. No question names a clinic.
- Engines
- ChatGPT with browsing (gpt-5-search-api), ChatGPT without browsing (gpt-5.6-terra), Perplexity (sonar), Gemini (gemini-3.6-flash), and Google AI Overview read from the live search page, pinned to each city.
- Repeats
- Three full passes of everything, so the engines' own instability can be measured instead of assumed.
- Answers
- 1,800 collected, zero errors — 600 per market. Usable answers per language: Kyiv 281 / 281, Warsaw 252 / 252, Vilnius 275 / 252. Percentages divide by each language's own figure.
- Dates
- Kyiv 22 August 2026, Warsaw 23 August, Vilnius 24 August. Each market's six passes ran inside a single session.
- Counting
- Aggregate percentages only. A clinic counts once per answer that names it. Spelling, spacing and — in Vilnius — translated forms of one clinic are merged; that merging is measured, not assumed.
- Access
- Provider APIs, not the consumer apps. Answers may differ from what a person sees in the ChatGPT window.
We publish no ranking positions. What we are not claiming is a section, not a footnote. The full brand-mention matrices are downloadable below.
The finding
Take the set of clinics an engine names in the local language, and the set it names in English, and measure how much they overlap. Then — this is the part most comparisons skip — measure how much a language's own brand set overlaps with itself when you simply ask again an hour later. The second number is the noise floor. Without it, the first number means nothing.
That second number is never 100%, and what it is made of matters. Ask Vilnius the same
twenty questions in Lithuanian at 09:30 and again at 11:00 and the engines name
Tavo dantistas, SSKC and Implantų centras the first time
and not the second, while Fi Clinica and Neodenta appear only the
second time. Nothing was rephrased and nothing was translated. Those are real clinics, dropping
in and out of the recommendations of an engine asked an identical question in an identical
language ninety minutes apart. About 15% of the cast churns that way, and any honest reading of
this category has to start there: a list of clinics that changes when you ask twice is
not evidence of anything until you know how much it changes when nothing changes.
That is what the noise floor buys. Cross-language overlap has to beat it to mean anything. We ran the comparison three times, in three countries, against three different populations of clinics. The same-language overlap is higher than the cross-language overlap in every one of them.
Do not read the three gaps against each other. The casts are different sizes — 48 clinics in Kyiv, 113 in Warsaw, 60 in Vilnius — and this measure falls as a cast grows, on both sides of the comparison. The noise floors move with it. What replicates is the direction, the separation from the noise floor and the per-engine pattern. A claim that the effect is strongest in Poland would need one cast-construction rule applied to all three markets, and we do not have that.
The engine that does not search is the engine that does not care what language you use
Splitting the gap by engine is where the mechanism shows. Four of the five engines answer every question asked; the fifth, Google's AI Overview, often declines to answer at all, so it is handled separately below.
ChatGPT without browsing shows no language effect in any of the three markets — +1, +1 and +4, with all three ranges crossing zero. Fewer than half the resamples came out positive at all (45%, 42%, 43%), which is a coin flip: these are not small effects, they are an absence of evidence for any effect. The three engines that retrieve documents before answering show the effect in all three. Three languages, three clinic populations, three independent replications of the same null.
That is the strongest claim this study supports: the language effect travels with retrieval. A model answering from its own weights gives you roughly the same clinics whichever language you use. A model that goes and reads the web first gives you different ones, because the web it reads is different in each language.
What this does not license is a league table of engines. Neighbouring intervals overlap heavily. Perplexity comes out top in all three markets, which is suggestive and is not the same as demonstrated.
The leaders are safe. The reshuffle happens below them.
Raising the bar for how often a clinic has to be named before it counts collapses the gap in every market — and it collapses to nothing at the top.
| Clinics counted | Kyiv | Warsaw | Vilnius |
|---|---|---|---|
| All of them | +16 | +27 | +12 |
| Named 10+ times | +9 | +12 | +4 |
| Named 20+ times | +4 | +0 | +4 |
| Named 40+ times | +2 | +0 | +0 |
Among the clinics an engine names most often, the two languages agree almost completely. In Vilnius the two biggest names are level to within half a point — 27.3% against 27.0%, and 26.5% against 29.0%. The entire effect lives below that line.
This is the commercially useful half of the finding, and it is the same shape in all three countries: if you are the category leader you are safe in either language. If you are anyone else, the language of the question reshuffles you.
Google sometimes answers the question in one language and refuses it in the other
Google's AI Overview does not appear on every search. Whether it appears at all turned out to be the sharpest version of this study's thesis — and the reason that engine is excluded from the chart above.
| Overview served on… | Local language | English |
|---|---|---|
| Kyiv | 68% of questions | 68% |
| Warsaw | 20% | 20% |
| Vilnius | 58% | 20% |
In Vilnius, asking in English rather than Lithuanian means Google declines to answer at all roughly two-thirds of the times it otherwise would. Not which clinics it names — whether it responds at all.
The decision is also oddly mechanical. In Warsaw and in English-language Vilnius, the same questions got an overview on all three passes. In Kyiv the set churned from pass to pass. Whatever makes Google withhold an overview looks deterministic per query in some markets and weather-like in others.
We do not publish a cross-language gap for this engine. In Vilnius the two languages share only three of the questions Google answers, so that comparison would largely be a comparison of different questions. The withholding rate is the finding here; the brand overlap is not measurable cleanly enough to report.
How hard we tried to break this
Every check below was run after the data was collected and is published whichever way it came out. The full outputs are in the repository alongside the raw answers.
The curation is the soft spot, and each market needed a different rule
Before you can count how often a clinic is named, you have to decide what counts as the same clinic — that Люмі-Дент and Lumi-Dent are one business, that a clear-aligner brand is not a clinic at all. Those are judgement calls, we made them, and they change the number. So the claim in this section is not that we avoided them. It is that we measured what each one was worth and published the reading least flattering to our own thesis. Each country forced a different call:
- Kyiv — alphabet. The engines write the same clinic as Люмі-Дент in Ukrainian and Lumi-Dent in English. Left unmerged, a change of script would have counted as a change of recommendation.
- Warsaw — nothing. Polish shares an alphabet with English and leaves clinic names in the nominative, so no judgement call was needed at all. That makes Warsaw the cleanest of the three tests.
- Vilnius — translation. Lithuanian clinic names are often descriptive, and the engines translate them: Dantų harmonija becomes "Dental Harmony", Vilniaus implantologijos centras becomes "Vilnius Implantology Center". Left unmerged, each becomes a brand that exists only in English — this study's headline claim, manufactured out of a translation.
Concretely. The Lithuanian answers name Dantų harmonija. The English answers,
about the same business on the same street, call it “Dental Harmony”. Rule that those are one
clinic and the two languages agree about it. Rule that they are two and the Lithuanian
list now holds a clinic the English list never mentions, while the English list holds one that
appears nowhere in Lithuanian — two disagreements conjured out of one translated name. The
second rule widens the exact gap this study exists to report.
So we tallied Vilnius both ways, over the same 600 answers. Every column is an overlap: of all the clinics named across two runs, the share named in both.
| How we counted clinic names | Lithuanian vs Lithuanian (noise floor) | Lithuanian vs English | Gap |
|---|---|---|---|
| One clinic — published | 85% | 73% | +12 |
| Two different clinics | 85% | 68% | +16 |
The noise floor does not move — 85% in both rows — and that is the whole mechanism. The English gloss never appears in the Lithuanian answers, so how we count it cannot affect a Lithuanian-to-Lithuanian comparison. It can only change the cross-language number. Left unmerged, the engine's own translation of a name shows up as a clinic that exists in English and not in Lithuanian, which is precisely the finding this study claims to have measured.
We publish the smaller number. A third of the unmerged gap would have been an artefact of an engine translating a name. Kyiv's own curation was worth about half of its headline by the same kind of test, which is why that market's honest claim is a range rather than a point.
Traps that would have manufactured the finding
Candidate brand names are mined out of the answers rather than guessed in advance — the whole point is that we do not know who the engines will name. That process throws up things that look like clinics and are not, and they do not land evenly across languages:
- Ordoline is a clear-aligner brand, listed in the answers beside Invisalign. It reads exactly like a clinic name.
- Pincetas is a review platform, cited beside Google Maps — despite being named in 21 separate answers.
- vilniusdental.lt is a real clinic and is still dropped, because in English its name is not separable from the ordinary phrase "Vilnius dental clinics". Nine of its 28 English matches came only from that generic phrasing, against none of its 13 Lithuanian ones. Keeping it would have added nine English-only mentions and moved the gap our way.
- In Warsaw the same rule removed Modern Dental ("modern dental care"), Right Clinic ("the right clinic") and Dental Bridge — a procedure, when one of the twenty questions asks about implants versus bridges.
Four more checks
- Unequal denominators. Vilnius is the one market where the two languages did not yield the same number of usable answers — 275 against 252, entirely because Google withheld more English overviews. More answers means more chances to name a clinic, which can widen the gap by itself. Dropping that engine equalises the counts at 240 and 240 and moves the gap from +12 to +11. The asymmetry is worth one point.
- Citations are not driving it. A clinic whose domain merely appears in a footnote is not a clinic the engine recommended. Stripping URLs before matching leaves 2.9% of Lithuanian and 2.7% of English mentions URL-only — symmetric, so it cannot produce a cross-language gap. Warsaw: 1.4% and 1.1%.
- The cast was not revised to fit the hypothesis. In Warsaw and Vilnius the brand list was mined from all six passes at once and curated before any overlap number was computed. Kyiv's was extended after the pattern was already visible — the right read of the data and the wrong procedure — and that market's writeup says so.
- Is the language split special? There are ten ways to divide six runs into two groups of three. In every market the true language split ranks first, and the next best split scores near zero. With six runs this test cannot produce a p below 0.10, and we report that ceiling rather than hiding it.
What we are not claiming
- Not a ranking. We publish no "position in AI". Repeat runs of the same prompt return different lists; anything calling itself a rank is measuring noise.
- Not a day-scale noise floor. Each market's six passes ran inside one session. Instability over days or weeks could be larger than what we measured, and if it is, these gaps narrow.
- Not a magnitude that travels. The three gaps are not comparable to each other, for the reason given above.
- Not a per-clinic result. Individual clinics are shown to illustrate the aggregate, not as findings in their own right. Dozens of clinics were tested per market with no correction for multiple comparisons, so some individual rows are chance.
- Not the consumer apps. These are provider APIs. What a person sees in the ChatGPT window may differ.
- Not a claim about any named business. We name the clinics engines do recommend. We do not publish a list of businesses that are invisible.
The data
Every number above can be recomputed from these files. Each row is a clinic, in one language, on one pass, from one engine, with the number of answers naming it and the denominator of answers that engine returned.
- kyiv-brand-matrix.csv — 1,440 rows
- warsaw-brand-matrix.csv — 3,390 rows
- vilnius-brand-matrix.csv — 1,800 rows
Published under CC BY 4.0. If you use it, a link back is enough.
Why we ran this
QuotedIn tracks how AI answer engines mention a business against its competitors, and reports it weekly. Every tool in this category is English-first, which is precisely the gap this study measures. If you run marketing for clients outside the English-speaking world, a free one-off scan will tell you what these engines say about them — in both languages.