Google Search Console for Social Media: How to Analyze Instagram, TikTok, and YouTube Visibility in Search
Transforming Text Content into Digital Assets for AI Search
5 Strategies for neutralizing undesirable online information
According to recent research, ChatGPT shows a clear preference for English-language pages, unlike Google AI and Copilot. Below is an analysis of what this means for your website.
If there is a consistent trend in which English-language pages are cited in AI-generated responses more frequently than equivalent pages in local languages, this assessment can be considered objective.
To evaluate this phenomenon, data from Copilot, Google AI, and ChatGPT were analyzed across dozens of multilingual websites. The goal was to measure the degree of language bias in these AI systems and determine whether adding an English-language version (using an /en/ directory) is worthwhile.
Website owners and operators outside the English-speaking market have repeatedly observed the same pattern: when users submit prompts in Dutch, French, Spanish, or other languages, AI systems often rely primarily on English-language sources.
When a multilingual website includes an English version, its English pages are more likely to be cited by AI models, even if they receive relatively little organic traffic. SEO experts attribute this behavior to one key factor: before retrieving information, AI systems often transform a non-English prompt into multiple English search requests, known as fan-out queries.
In The AI Visibility Audit (Common Crawl), Stephen Burns notes: “A high-quality English version of your site’s key pages is one of the few content investments that genuinely expands a brand’s digital visibility.”
Given the growing importance of this topic for our agency’s clients, we conducted our own practical analysis centered around two key questions:
The strongest evidence of this preference comes from the behavior of OpenAI’s web crawlers, based on the analysis of server logs. OpenAI operates three different crawlers, and treating them as a single mechanism can lead to an inaccurate understanding of how the system works:
The final list of URLs cited in ChatGPT responses is determined primarily by the activities of OAI-SearchBot and ChatGPT-User.
Log data segmented by crawler shows that the preference for English is not static. Instead, it increases significantly as the system moves closer to the citation and response-generation stage.
A comparison of crawler behavior reveals several notable patterns. GPTBot, the crawler responsible for gathering training data, shows little to no preference for English-language pages on multilingual websites. It crawls content in all available languages, and the share of English-language content it visits is actually lower than the level of real organic demand for English.
From a model training perspective, this is expected. Since the foundation models were initially trained on a large volume of English-language data, OpenAI benefits from continuously collecting content in other languages to improve multilingual capabilities.
OAI-SearchBot demonstrates a moderate level of English-language preference.
The strongest bias appears in ChatGPT-User—the crawler whose activity occurs immediately before a source is cited. It retrieves English-language versions of pages more frequently than any other OpenAI crawler. Its median rate of English-language retrieval is approximately three times higher than that of Googlebot on the same multilingual websites.
Therefore, the common statement that “ChatGPT prefers English” requires an important clarification. The preference is primarily exhibited by the retrieval crawler, not by the training crawler. This distinction is significant because it suggests that the bias emerges during the information retrieval stage rather than during the model training process itself.
Although most of our agency’s clients are based in Europe—including Belgium, the Netherlands, and France—the same pattern has also been observed on websites in Canada and India.
The level of overrepresentation (or over-citation) of English-language pages ranges from 1.9× to 6.5×, with a median value of approximately 2.6×. Overall, ChatGPT-User retrieves English-language pages in roughly 65–79% of cases.
These findings indicate a consistent pattern: before generating its final response, the AI assistant first determines which language version of a website or brand to analyze, with a clear tendency to prioritize English-language content.
The way prompts are processed varies significantly depending on the AI model. When the same analysis of English-language overrepresentation was conducted for Copilot (using the AI Performance report in Bing Webmaster Tools) and for Google AI Mode and AI Overviews (based on the latest reports in Google Search Console), the observed bias was found to be minimal or nonexistent.
ChatGPT clearly favors English-language content, while Copilot, Google AI Mode, and AI Overviews do not.
To quantify this effect, a single metric was calculated: the ratio between the share of English-language citations generated by an AI system and the share of impressions that English-language content receives in traditional search results. A value of 1.0 represents a neutral baseline, meaning that the AI cites English-language content at the same rate that the website receives impressions from English-language search queries. Values above 1.0 indicate a disproportionate preference for English.
Copilot records a score of 1.07 (based on a sample of 272 websites in Bing), which is effectively neutral.
Google AI, according to Google Search Console data, scores 0.79, meaning it references slightly less English-language content than would be expected based on search demand, instead favoring content in the user’s local language.
By comparison, ChatGPT reaches an index of approximately 2.6.
The reasons for these differences lie in the underlying architecture of each system. Copilot and Google AI rely on search indexes that have been optimized for multilingual retrieval over many years. As a result, they inherit mature mechanisms for delivering language-appropriate content to users, including support for hreflang tags, and their AI-generated responses largely preserve this logic.
ChatGPT, by contrast, performs real-time information retrieval within a model whose primary knowledge foundation is English-language content, without relying on the same mature multilingual search infrastructure. Consequently, the observed bias appears to be a characteristic of the LLM retrieval path rather than of search indexing itself.
For SEO professionals and website owners, this distinction provides an important insight: the phenomenon described here is specific to ChatGPT and is likely shared by other large language models that use a retrieval architecture similar to OpenAI’s.
Google once introduced the concept of mobile-first indexing, making the mobile version of a website the primary source for indexing and ranking.
For large language models (LLMs), particularly ChatGPT, English-language content appears to be taking on a similar role. When a website includes an English version, some AI systems may treat it as the primary source of information about the site, regardless of how important that version is to the website’s actual target audience.
This emerging English-first paradigm creates potential risks for websites where the English section was originally intended as a secondary or supporting version. In one case identified during our analysis, a financial institution had left an outdated English-language PDF containing pricing information on its website, while the Dutch and French versions had already been updated.
The AI assistant retrieved and cited the outdated English document instead of the current local-language pages. As a result, an obsolete English-language file became the final answer presented to users—even though they had not submitted their query in English.
Since large language models tend to prioritize English-language sources, providing them with high-quality English content is a logical strategy. To evaluate its effectiveness, we compared AI visibility with visibility in traditional search.
Using a 1:1 baseline between traditional search visibility (Google and Bing impressions) and AI visibility, we compared websites with and without an /en/ section. The results revealed a clear pattern.
In Copilot, websites with an English-language section received 892 citations per 10,000 search impressions, compared with 585 citations for websites without one—a difference of 52%.
To validate this finding, the same analysis was performed using Google data. According to the Google Search Console Gen AI report, websites with an English version achieved 28% higher AI visibility overall. After removing statistical outliers, the median increase per website was 9%, further highlighting the differences between AI platforms.
Data from OpenAI’s crawlers confirms that the strongest preference for English-language content is observed in ChatGPT. Websites with an /en/ section experienced an increase in AI visibility of up to 122%.
Overall, the findings support Stephen Burns’s thesis. For most of the multilingual websites analyzed, the presence of a high-quality and up-to-date English version of key (cornerstone) pages contributed to greater visibility in AI systems. The effect was strongest in ChatGPT, noticeable in Copilot, and directionally similar across Google’s AI-powered search surfaces.
However, two important factors prevent a universal recommendation that every website should add an /en/ section.
First, maintaining an English-language version is an ongoing process, not a one-time task. Once published, large language models may treat the English version as canonical. An outdated English page can be more harmful than having no English version at all, as illustrated by the banking PDF example discussed earlier. An English section should therefore be introduced only if the organization can maintain it with the same update frequency as the primary-language content.
Second, the advantage of English-language content appears less pronounced in more mature AI systems, while OpenAI continues to expand training on non-English data. This suggests that the incremental visibility benefit of English-only content may decline over time.
As a result, creating an English-language section is justified only when sufficient resources are available for its long-term maintenance and continuous updating.
The study is based on three independent datasets:
For the purposes of this study, English-language content refers to local URL paths matching patterns such as /en, /en-xx, /xx-en, or a root domain dedicated exclusively to English-language content.
The representation index is calculated as the ratio of the share of English-language AI citations to the share of English-language impressions in traditional search, where a value of 1.0 represents a neutral baseline.
English-language request shares for each crawler were calculated using the median values across websites that recorded at least 150–200 requests per crawler, reducing distortion from unusually large domains. Overrepresentation ratios compare each crawler’s English-language share with the Googlebot English-language share on the same website.
AI citation reporting is still evolving and contains statistical noise. Bing and Google Search Console use different methodologies for defining AI impressions and citations; therefore, cross-platform comparisons should be interpreted as directional rather than as directly proportional absolute values.
Crawler identification in server logs relies on User-Agent strings and reverse DNS verification, which may fail to detect spoofed or proxied traffic.
Several parts of the analysis, including the crawler-level breakdown and portions of the GSC panel, are based on relatively small website samples ranging from a few sites to several dozen. These results should therefore be interpreted as indicative rather than definitive. Brand names and domains were anonymized according to their industry categories.
Read this article in Ukrainian.
Say hello to us!
A leading global agency in Clutch's top-15, we've been mastering the digital space since 2004. With 9000+ projects delivered in 65 countries, our expertise is unparalleled.
Let's conquer challenges together!
performance_marketing_engineers/
performance_marketing_engineers/
performance_marketing_engineers/
performance_marketing_engineers/
performance_marketing_engineers/
performance_marketing_engineers/
performance_marketing_engineers/
performance_marketing_engineers/