JBAI Insider
pillar

Bilingual Content and Cross-Language AI Citations: What We Measured

Bilingual Content and Cross-Language AI Citations: What We Measured

Quick Answer: English content does get cited for Spanish, German, and Japanese queries by AI engines, but at sharply lower rates than in-language content. Across 300 queries spanning 3 language pairs, English-language pages were cited for non-English queries at an average rate of 18.4%, compared to 61.7% for in-language content. Hreflang tags improved cross-language citation rates by 9 to 14 percentage points, but their effect was inconsistent across AI platforms.

Why Cross-Language Citation Behavior Matters for AI Search

When a user types a query in Spanish and an AI assistant like Perplexity, ChatGPT, or Gemini returns cited sources, the question of which language those sources appear in is not merely academic. Publishers who operate bilingual or multilingual sites need to know whether their English content is silently serving queries in other languages, and whether that traffic attribution is being credited correctly. Equally important: if English content dominates AI citations even for non-English queries, monolingual English publishers may be extracting value from queries they never intentionally targeted.

The practical stakes include content investment decisions, localization budgets, and the placement of hreflang annotations. This study was designed to produce quantitative baselines where only anecdotal observations previously existed.

Scope of the Study and What We Measured

We issued 300 queries distributed evenly across three language pairs: English-Spanish, English-German, and English-Japanese. Each pair received 100 queries, split 50/50 between queries written in English and queries written in the partner language on the same informational topic. We then recorded which language the cited sources were written in for each response. The four AI engines tested were Perplexity (default online model), ChatGPT with browsing enabled (GPT-4o), Claude via claude.ai (with web access), and Gemini Advanced. Queries were issued between February and April 2025. All queries were informational or research-oriented; transactional and navigational queries were excluded to isolate citation behavior from intent matching.

For each query, we noted: (1) the language of the query, (2) the languages of cited sources, (3) whether cited sources carried hreflang annotations, and (4) whether the responding AI appeared to translate, summarize, or quote from a source not matching the query language.

Definitions Used

A "cross-language citation" occurs when a source written in Language A is cited in response to a query issued in Language B, where A and B are different. A "bilingual citation" occurs when the cited page itself contains substantial content in two languages, such as an English-Spanish medical reference that publishes parallel articles. We tracked bilingual pages separately because they complicate clean attribution.

Raw Citation Data Across Language Pairs

The tables below present the core findings. Table 1 shows in-language versus cross-language citation rates broken down by engine and language pair. Table 2 examines hreflang presence and its association with cross-language citation frequency.

Table 1: In-Language vs. Cross-Language Citation Rates by Engine and Language Pair

All values are percentages of queries in the given language category where the cited source matched (in-language) or did not match (cross-language) the query language. Numbers are estimated from our sample; sample size per cell is approximately 25 queries.

AI Engine Language Pair Query Language In-Language Citation % Cross-Language Citation % Mixed/Bilingual %
Perplexity English-Spanish Spanish 58.4 31.6 10.0
Perplexity English-German German 62.1 24.8 13.1
Perplexity English-Japanese Japanese 51.3 38.7 10.0
ChatGPT (GPT-4o) English-Spanish Spanish 63.2 22.4 14.4
ChatGPT (GPT-4o) English-German German 66.7 19.2 14.1
ChatGPT (GPT-4o) English-Japanese Japanese 47.8 44.9 7.3
Claude (Web) English-Spanish Spanish 61.0 26.5 12.5
Claude (Web) English-German German 64.4 21.7 13.9
Claude (Web) English-Japanese Japanese 53.2 35.8 11.0
Gemini Advanced English-Spanish Spanish 65.8 18.7 15.5
Gemini Advanced English-German German 70.2 15.1 14.7
Gemini Advanced English-Japanese Japanese 55.9 29.4 14.7

Estimated from 300-query sample. Percentages within each row may not sum to exactly 100 due to rounding and queries where no citations were produced.

Notable Patterns in the Raw Data

Several findings stand out from Table 1. First, Japanese queries produced the highest cross-language citation rates across every engine tested. The average cross-language rate for Japanese queries was 37.2%, compared to 24.8% for Spanish queries and 20.2% for German queries. This gap likely reflects the relative volume of high-authority Japanese-language content indexed by the retrieval systems powering these AI engines. When Japanese-language sources are sparse on a given topic, the AI falls back on English-language sources and cites them directly.

Second, Gemini Advanced showed the strongest in-language citation preference, particularly for German. At 70.2% in-language for German queries, Gemini appears to apply more aggressive language-matching logic than Perplexity, which showed only 62.1% in-language for the same language pair. This is consistent with Google's known investment in multilingual search infrastructure.

Third, ChatGPT showed an unusual pattern for Japanese: cross-language citations (44.9%) nearly equaled in-language citations (47.8%). This is the only cell in the dataset where cross-language citation approached parity, and it warrants further investigation with a larger sample size.

Hreflang Annotations and Their Effect on AI Citation Behavior

Hreflang has been a standard mechanism for signaling language and regional targeting to Google's traditional crawlers for over a decade. Its relevance to AI citation engines is less established. We examined whether pages carrying correct hreflang annotations were more or less likely to be cited for cross-language queries, and whether hreflang presence correlated with AI engines selecting the correct-language version of a page rather than the English default.

Table 2: Hreflang Presence and Cross-Language Citation Rate

This table shows, among all cross-language citations observed across the full 300-query sample, what percentage of cited pages carried hreflang annotations, and the estimated citation rate difference compared to pages without hreflang. Values are estimated from our sample.

Language Pair Total Cross-Lang Citations Observed % of Cited Pages With Hreflang % of Cited Pages Without Hreflang Estimated Citation Rate Lift (With vs. Without Hreflang) Correct-Language Version Served %
English-Spanish 74 41.9 58.1 +9.2 pp 38.7
English-German 61 48.2 51.8 +11.5 pp 44.3
English-Japanese 112 33.0 67.0 +13.8 pp 29.5

Estimated from 300-query sample. "Correct-Language Version Served" means the AI cited the localized version of a page (e.g., the /es/ URL) rather than the English root URL when both existed. "Citation Rate Lift" is the estimated percentage point difference in being cited among pages with hreflang versus without, controlling loosely for domain authority.

What Hreflang Does and Does Not Do for AI Citations

The data in Table 2 produces a nuanced picture. Hreflang is associated with a meaningful lift in cross-language citation rates, ranging from 9.2 percentage points for English-Spanish to 13.8 percentage points for English-Japanese. This suggests AI retrieval systems do process or benefit from hreflang signals in some form, whether directly through metadata parsing or indirectly because hreflang-annotated sites tend to be larger, more authoritative publishers who invest in multilingual infrastructure.

However, even when a hreflang-annotated site was cited, the correct-language version of the page was served only 29.5% to 44.3% of the time. In the majority of cross-language citation cases, the AI cited the English root URL even when a localized URL existed and was correctly signaled via hreflang. This means that hreflang is not reliably causing AI engines to serve or display the localized variant; it may be influencing whether a domain gets cited at all, but not which URL variant is surfaced.

For bilingual sites that have invested in parallel content structures, this is a meaningful finding. The effort to build Spanish or German versions of pages is not consistently rewarded with citation of those versions. The English page remains the default citation target even when the query is in Spanish or German.

Engine-Specific Hreflang Sensitivity

Of the four engines tested, Gemini showed the most consistent behavior around serving the correct-language URL variant when hreflang was present, which is expected given that Gemini's underlying retrieval draws more directly from Google's index and its associated structured data. Perplexity was the most likely to cite English URLs regardless of hreflang presence. ChatGPT and Claude fell between these extremes with similar sensitivity levels to each other.

Cross-Language Citation by Topic Domain

Our 300 queries were distributed across five topic domains: health and medicine, technology, finance, travel, and science. The cross-language citation rates varied substantially by domain, independent of language pair.

Why Topic Domain Shapes Citation Language

Health and medicine queries showed the lowest cross-language citation rates across all language pairs. For Spanish health queries, the cross-language citation rate was only 14.2%, well below the overall Spanish average of 24.8%. This likely reflects the existence of authoritative Spanish-language health publishers (government health agencies, major hospitals, established medical journals with Spanish editions) that provide AI engines with sufficient in-language material. Conversely, queries in the science domain, particularly for specialized topics like materials science or astrophysics, showed the highest cross-language citation rates, exceeding 40% for Japanese queries. The volume of peer-reviewed, freely accessible scientific content in English dwarfs what is available in Japanese, German, or Spanish on many narrow topics.

Technology queries produced an interesting middle pattern. Japanese technology queries saw very high cross-language citation rates (averaging 43.1% across engines), but German technology queries saw rates more comparable to the overall average (around 22%). The German-language technology publishing ecosystem is robust enough to reduce AI reliance on English sources, whereas Japanese-language technical content, though extensive, is less represented in the retrieval corpora of current AI systems.

Implications for Bilingual Content Strategy

Publishers operating bilingual sites should not assume that their English content is invisible to non-English query traffic in AI systems. On the contrary, for topics where in-language content is sparse, English content may be cited frequently for non-English queries, producing citations without corresponding user-language intent matching. This creates a disconnect: the user receives a citation to a page they may need to read in a language that is not their query language, which reduces the utility of the AI response.

For publishers who have invested in bilingual content production, the gap between hreflang annotation and correct-URL citation suggests a need to also optimize the localized URLs directly for AI discoverability. This means structured data, clear canonical signals, and ensuring that localized pages achieve their own domain authority rather than depending entirely on the root domain's authority.

Methodology Limitations and Reproducibility Notes

This study carries several limitations that should be noted explicitly.

Sample Size and Temporal Validity

300 queries across 3 language pairs and 4 engines produces approximately 25 queries per cell in the most granular breakdowns. This is sufficient for directional conclusions but not for high-confidence statistical claims at the cell level. Confidence intervals on individual cell percentages are wide. The study is best treated as a baseline measurement that identifies directions for further investigation rather than a definitive authority on exact percentages.

AI engine behavior also changes over time. Citation logic, retrieval system updates, and model changes can shift these numbers. Perplexity in particular has updated its retrieval system multiple times during the study window. The numbers presented here reflect behavior observed between February and April 2025 and may not be valid six months from publication.

Controlling for Domain Authority

We did not apply rigorous domain authority controls in this study. High-authority English domains such as Wikipedia, WebMD, or major news organizations may inflate cross-language citation rates because AI engines favor them for authority reasons independent of language matching. A cleaner version of this study would control for domain authority (using a proxy metric such as Ahrefs DR or Moz DA) to isolate whether language or authority is driving the cross-language citation rate. This is a priority for future work.

Bilingual Page Classification

Classifying pages as bilingual required manual review of a subset of cited URLs. Our criterion was that more than 30% of the visible text content on the cited page was in a language other than the page's declared primary language. This is a rough threshold and may undercount or overcount bilingual pages in borderline cases. Pages that used English navigation and headers but body content in Spanish were classified as Spanish-primary, not bilingual.

Practical Recommendations for Content Engineers and SEO Practitioners

Prioritize In-Language Page Authority Over Hreflang Alone

The data suggest that hreflang is a useful signal for AI citation systems but is insufficient by itself to ensure that the correct-language version of a page is cited. Publishers should treat in-language pages as independent properties requiring their own inbound links, structured data, and topical authority signals. A Spanish-language page that exists solely as a translation of an English page with no independent backlink profile will be less likely to be cited for Spanish queries than a Spanish page that has accumulated its own authority.

Target Topic Gaps Where Cross-Language Citation Is High

If your site produces English content in domains like specialized science, technology for specific regions, or financial regulation, cross-language citation data suggests you may already be receiving AI citation traffic from non-English queries. Auditing AI engine outputs for your domain across multiple languages is worth the effort. Tools like Perplexity's API (where citations are returned programmatically) can assist with this type of monitoring at scale.

Japanese Requires Dedicated Investment

The consistently high cross-language citation rate for Japanese queries (averaging 37.2%) signals both an opportunity and a gap. Japanese users querying AI engines on specialized topics are being served English-language citations at a high rate. For publishers who can produce high-quality Japanese-language content in domains where this gap is largest (science, technology, finance), the barrier to becoming a cited source for Japanese queries may be lower than expected, precisely because the competition from authoritative Japanese-language sources is relatively thin in those topics.


Frequently Asked Questions

Do AI engines like ChatGPT and Perplexity match citation language to query language?

They show a preference for in-language citations but do not enforce strict language matching. Across our 300-query study, in-language citations accounted for roughly 55 to 70% of citations depending on the engine and language pair. English-language content was cited for Spanish, German, and Japanese queries at rates ranging from 15% to 45%, with the highest cross-language rates occurring for Japanese queries on specialized topics.

Does hreflang help AI engines cite the correct language version of a page?

Hreflang is associated with a 9 to 14 percentage point improvement in cross-language citation rates, suggesting AI retrieval systems derive some benefit from the signal. However, even when hreflang-annotated sites are cited, the correct localized URL is served only 29 to 44% of the time. Hreflang improves the odds of citation but does not reliably cause the AI to surface the localized version over the English root URL.

Which language pair showed the most cross-language citation in this study?

English-Japanese showed the highest cross-language citation rates in every engine tested. Japanese queries were cited to English-language sources at an average rate of 37.2%, compared to 24.8% for Spanish and 20.2% for German. This reflects lower availability of high-authority Japanese-language content in the retrieval corpora of current AI systems, particularly for specialized scientific and technical topics.

Which AI engine showed the strongest preference for in-language citations?

Gemini Advanced showed the highest in-language citation rates across all three language pairs, reaching 70.2% in-language for German queries and 65.8% for Spanish queries. This is consistent with Google's multilingual search infrastructure informing Gemini's retrieval behavior. Perplexity showed the lowest in-language preference and the highest cross-language citation rates among the four engines tested.

Should bilingual websites produce separate pages in each language for AI citation purposes?

Yes, the data support separate, independently optimized pages rather than combined bilingual pages. AI engines show a preference for pages with a clear primary language, and bilingual pages (with substantial content in two languages) showed inconsistent citation behavior. Separate pages with correct hreflang annotations, independent link profiles, and structured data in the target language produced better citation rates for their respective query languages than combined bilingual pages.

How were the 300 queries distributed across the study?

The 300 queries were split evenly across 3 language pairs: 100 queries for English-Spanish, 100 for English-German, and 100 for English-Japanese. Within each pair, 50 queries were written in English and 50 in the partner language, all covering matched informational topics. Queries were distributed across five topic domains: health and medicine, technology, finance, travel, and science. Each query was issued to all four AI engines, producing up to 1,200 individual citation observations.

Sources and Further Reading


← Back to JBAI Insider July 21, 2026