QuotedIn AI visibility reports for agencies

Study · Kyiv · Warsaw · Vilnius · August 2026

Does the language you ask in change which businesses AI recommends?

We asked five AI answer engines the same twenty dental questions about three capital cities — once in the local language, once in English — and repeated the whole thing three times in each. 1,800 answers. The recommended clinics differ by language in all three markets, beyond what the engines' own instability explains. The effect is concentrated below the top of the leaderboard, and how big it looks depends on how you count.

Method, before the numbers

Markets
Kyiv (Ukrainian / English), Warsaw (Polish / English), Vilnius (Lithuanian / English). Three countries, so the result reads as a property of answer engines rather than a quirk of one language.
Questions
20 buying questions a real patient would type, written natively in each language and translated 1:1. The same twenty questions in every market. No question names a clinic.
Engines
ChatGPT with browsing (gpt-5-search-api), ChatGPT without browsing (gpt-5.6-terra), Perplexity (sonar), Gemini (gemini-3.6-flash), and Google AI Overview read from the live search page, pinned to each city.
Repeats
Three full passes of everything, so the engines' own instability can be measured instead of assumed.
Answers
1,800 collected, zero errors — 600 per market. Usable answers per language: Kyiv 281 / 281, Warsaw 252 / 252, Vilnius 275 / 252. Percentages divide by each language's own figure.
Dates
Kyiv 22 August 2026, Warsaw 23 August, Vilnius 24 August. Each market's six passes ran inside a single session.
Counting
Aggregate percentages only. A clinic counts once per answer that names it. Spelling, spacing and — in Vilnius — translated forms of one clinic are merged; that merging is measured, not assumed.
Access
Provider APIs, not the consumer apps. Answers may differ from what a person sees in the ChatGPT window.

We publish no ranking positions. What we are not claiming is a section, not a footnote. The full brand-mention matrices are downloadable below.

The finding

Take the set of clinics an engine names in the local language, and the set it names in English, and measure how much they overlap. Then — this is the part most comparisons skip — measure how much a language's own brand set overlaps with itself when you simply ask again an hour later. The second number is the noise floor. Without it, the first number means nothing.

That second number is never 100%, and what it is made of matters. Ask Vilnius the same twenty questions in Lithuanian at 09:30 and again at 11:00 and the engines name Tavo dantistas, SSKC and Implantų centras the first time and not the second, while Fi Clinica and Neodenta appear only the second time. Nothing was rephrased and nothing was translated. Those are real clinics, dropping in and out of the recommendations of an engine asked an identical question in an identical language ninety minutes apart. About 15% of the cast churns that way, and any honest reading of this category has to start there: a list of clinics that changes when you ask twice is not evidence of anything until you know how much it changes when nothing changes.

That is what the noise floor buys. Cross-language overlap has to beat it to mean anything. We ran the comparison three times, in three countries, against three different populations of clinics. The same-language overlap is higher than the cross-language overlap in every one of them.

3 of 3
Markets where the effect appears
Ukraine · Poland · Lithuania
100%
Resamples that still show the gap
2,000 per market, in all three
+12 … +27
Point gap over the noise floor
not comparable between markets — see below
Same language, different pass (noise floor) Local language vs English
Brand-set overlap within and across languages, three markets In every market the same-language overlap is higher than the cross-language overlap. Kyiv 87 versus 70 percent, Warsaw 78 versus 52, Vilnius 85 versus 73. Bars are measured from zero. 0%20%40%60%80%100% Kyiv same language 87% Kyiv: same language, different pass — 87% overlap uk vs en 70% Kyiv: Ukrainian vs English — 70% overlap Warsaw same language 78% Warsaw: same language, different pass — 78% overlap pl vs en 52% Warsaw: Polish vs English — 52% overlap Vilnius same language 85% Vilnius: same language, different pass — 85% overlap lt vs en 73% Vilnius: Lithuanian vs English — 73% overlap
Brand-set overlap (Jaccard) over every clinic the engines named in that market. Both bars are computed the same way — pairwise, between two runs — so the noise floor and the finding are directly comparable within a market. This is set overlap, which treats a clinic named once the same as one named two hundred times; weighted by how often each clinic is actually named, the gaps are +6 (Kyiv), +20 (Warsaw) and +18 (Vilnius).

Do not read the three gaps against each other. The casts are different sizes — 48 clinics in Kyiv, 113 in Warsaw, 60 in Vilnius — and this measure falls as a cast grows, on both sides of the comparison. The noise floors move with it. What replicates is the direction, the separation from the noise floor and the per-engine pattern. A claim that the effect is strongest in Poland would need one cast-construction rule applied to all three markets, and we do not have that.

The engine that does not search is the engine that does not care what language you use

Splitting the gap by engine is where the mechanism shows. Four of the five engines answer every question asked; the fifth, Google's AI Overview, often declines to answer at all, so it is handled separately below.

Whole range sits above zero — a real difference Range crosses zero — no difference shown
Language gap by engine and market, with 95% intervals The three retrieving engines show a positive gap in all three markets with intervals clear of zero. ChatGPT without browsing shows plus one, plus one and plus four, and all three intervals include zero. -20+0+20+40+60 Perplexity Kyiv +46 Perplexity, Kyiv: gap +46 points, 95% interval +37 to +54 Warsaw +53 Perplexity, Warsaw: gap +53 points, 95% interval +44 to +55 Vilnius +31 Perplexity, Vilnius: gap +31 points, 95% interval +27 to +47 Gemini Kyiv +20 Gemini, Kyiv: gap +20 points, 95% interval +15 to +31 Warsaw +29 Gemini, Warsaw: gap +29 points, 95% interval +18 to +35 Vilnius +23 Gemini, Vilnius: gap +23 points, 95% interval +8 to +30 ChatGPT, browsing Kyiv +10 ChatGPT, browsing, Kyiv: gap +10 points, 95% interval +1 to +19 Warsaw +13 ChatGPT, browsing, Warsaw: gap +13 points, 95% interval +8 to +17 Vilnius +10 ChatGPT, browsing, Vilnius: gap +10 points, 95% interval +3 to +17 ChatGPT, no browsing Kyiv +1 ChatGPT, no browsing, Kyiv: gap +1 points, 95% interval -13 to +12 — includes zero Warsaw +1 ChatGPT, no browsing, Warsaw: gap +1 points, 95% interval -13 to +12 — includes zero Vilnius +4 ChatGPT, no browsing, Vilnius: gap +4 points, 95% interval -12 to +9 — includes zero
Each dot is the gap, pooled across all twenty questions. Twenty is not many, and one unusual question can drag a pooled number a long way — so the line tests how much the result leans on any single one. We re-tallied the gap 2,000 times, each time from a random draw of our own twenty questions with some counted twice and some left out, using the answers already collected. No new questions were asked and no new API calls were made; it is the same data counted differently. The line covers the middle 95% of those 2,000 tallies. Where it stays entirely above zero, no single question is carrying the result. Where it crosses zero — the grey rows — dropping or repeating a couple of questions flips the sign, so nothing is shown. That is not the same as showing there is nothing: each engine answers only about twenty questions per pass, and a small effect would be invisible at that size.

ChatGPT without browsing shows no language effect in any of the three markets — +1, +1 and +4, with all three ranges crossing zero. Fewer than half the resamples came out positive at all (45%, 42%, 43%), which is a coin flip: these are not small effects, they are an absence of evidence for any effect. The three engines that retrieve documents before answering show the effect in all three. Three languages, three clinic populations, three independent replications of the same null.

That is the strongest claim this study supports: the language effect travels with retrieval. A model answering from its own weights gives you roughly the same clinics whichever language you use. A model that goes and reads the web first gives you different ones, because the web it reads is different in each language.

What this does not license is a league table of engines. Neighbouring intervals overlap heavily. Perplexity comes out top in all three markets, which is suggestive and is not the same as demonstrated.

The leaders are safe. The reshuffle happens below them.

Raising the bar for how often a clinic has to be named before it counts collapses the gap in every market — and it collapses to nothing at the top.

Clinics countedKyivWarsawVilnius
All of them+16+27+12
Named 10+ times+9+12+4
Named 20+ times+4+0+4
Named 40+ times+2+0+0

Among the clinics an engine names most often, the two languages agree almost completely. In Vilnius the two biggest names are level to within half a point — 27.3% against 27.0%, and 26.5% against 29.0%. The entire effect lives below that line.

This is the commercially useful half of the finding, and it is the same shape in all three countries: if you are the category leader you are safe in either language. If you are anyone else, the language of the question reshuffles you.

Google sometimes answers the question in one language and refuses it in the other

Google's AI Overview does not appear on every search. Whether it appears at all turned out to be the sharpest version of this study's thesis — and the reason that engine is excluded from the chart above.

Overview served on…Local languageEnglish
Kyiv68% of questions68%
Warsaw20%20%
Vilnius58%20%

In Vilnius, asking in English rather than Lithuanian means Google declines to answer at all roughly two-thirds of the times it otherwise would. Not which clinics it names — whether it responds at all.

The decision is also oddly mechanical. In Warsaw and in English-language Vilnius, the same questions got an overview on all three passes. In Kyiv the set churned from pass to pass. Whatever makes Google withhold an overview looks deterministic per query in some markets and weather-like in others.

We do not publish a cross-language gap for this engine. In Vilnius the two languages share only three of the questions Google answers, so that comparison would largely be a comparison of different questions. The withholding rate is the finding here; the brand overlap is not measurable cleanly enough to report.

How hard we tried to break this

Every check below was run after the data was collected and is published whichever way it came out. The full outputs are in the repository alongside the raw answers.

The curation is the soft spot, and each market needed a different rule

Before you can count how often a clinic is named, you have to decide what counts as the same clinic — that Люмі-Дент and Lumi-Dent are one business, that a clear-aligner brand is not a clinic at all. Those are judgement calls, we made them, and they change the number. So the claim in this section is not that we avoided them. It is that we measured what each one was worth and published the reading least flattering to our own thesis. Each country forced a different call:

Concretely. The Lithuanian answers name Dantų harmonija. The English answers, about the same business on the same street, call it “Dental Harmony”. Rule that those are one clinic and the two languages agree about it. Rule that they are two and the Lithuanian list now holds a clinic the English list never mentions, while the English list holds one that appears nowhere in Lithuanian — two disagreements conjured out of one translated name. The second rule widens the exact gap this study exists to report.

So we tallied Vilnius both ways, over the same 600 answers. Every column is an overlap: of all the clinics named across two runs, the share named in both.

How we counted clinic namesLithuanian vs Lithuanian (noise floor)Lithuanian vs EnglishGap
One clinic — published85%73%+12
Two different clinics85%68%+16

The noise floor does not move — 85% in both rows — and that is the whole mechanism. The English gloss never appears in the Lithuanian answers, so how we count it cannot affect a Lithuanian-to-Lithuanian comparison. It can only change the cross-language number. Left unmerged, the engine's own translation of a name shows up as a clinic that exists in English and not in Lithuanian, which is precisely the finding this study claims to have measured.

We publish the smaller number. A third of the unmerged gap would have been an artefact of an engine translating a name. Kyiv's own curation was worth about half of its headline by the same kind of test, which is why that market's honest claim is a range rather than a point.

Traps that would have manufactured the finding

Candidate brand names are mined out of the answers rather than guessed in advance — the whole point is that we do not know who the engines will name. That process throws up things that look like clinics and are not, and they do not land evenly across languages:

Four more checks

What we are not claiming

The data

Every number above can be recomputed from these files. Each row is a clinic, in one language, on one pass, from one engine, with the number of answers naming it and the denominator of answers that engine returned.

Published under CC BY 4.0. If you use it, a link back is enough.

Why we ran this

QuotedIn tracks how AI answer engines mention a business against its competitors, and reports it weekly. Every tool in this category is English-first, which is precisely the gap this study measures. If you run marketing for clients outside the English-speaking world, a free one-off scan will tell you what these engines say about them — in both languages.