A disquieting pattern has emerged in the behavior of leading artificial intelligence models: they exhibit a susceptibility to the information environments of authoritarian states, particularly China. For years, researchers and regulators have voiced concerns that Chinese-developed AI might inherently carry Beijing’s censorship. However, recent findings suggest that even Western-built models from companies like OpenAI, Anthropic, Google, and Meta are not entirely immune, raising questions about the supposed political neutrality these platforms often promise.
A peer-reviewed study published in *Nature* revealed how Chinese state-controlled media infiltrates AI training data, subtly influencing how these models respond to queries concerning China. The researchers identified over three million Chinese-language documents within the open-source training dataset CulturaX, a common resource for LLM development. Their analysis specifically focused on political subjects and found that models such as Claude Sonnet, Claude Opus, GPT-3.5 Instruct, GPT-4, and GPT-4o could reproduce distinctive phrases from Chinese state-coordinated media at rates ranging from 3% to nearly 10%. This suggests direct exposure during their training phases.
To further investigate this phenomenon, the *Nature* researchers conducted a targeted experiment. They took Meta’s open-weight Llama 2 13B model, which initially contained minimal Chinese state media in its training data, and then fine-tuned it with a small dataset of just 6,400 Chinese state-scripted news examples. The results were striking: the retrained model produced answers more favorable to Beijing almost 80% of the time compared to its baseline version. The disparity became even more pronounced with increased training data. When exposed to 64,000 state-scripted examples and asked if China is an autocracy, the baseline Llama 2 responded affirmatively. The state-influenced version, however, described China as democratic, echoing the Chinese Communist Party’s concept of “people’s democracy.”
Given the proprietary nature of systems from OpenAI and Anthropic, direct training data manipulation was not possible. Instead, researchers posed identical political questions in both Chinese and English to models like Claude Sonnet, Claude Opus, GPT-3.5, and GPT-4o. The responses in Chinese were rated as more favorable to Chinese leaders and institutions a significant portion of the time—68.8% for Claude Sonnet, 88.2% for Claude Opus, 72.6% for GPT-3.5, and 84% for GPT-4o. This language-dependent bias extended beyond China, with models offering more favorable descriptions of countries with lower press freedom when queried in their dominant language rather than in English. Researchers noted that this effectively “launders government-manipulated content into ostensibly objective text,” making it difficult for users to discern the true source or intent behind the information.
Further complicating the landscape, Meta’s independent Oversight Board uncovered a related issue: American AI models occasionally behave as if the political restrictions of authoritarian regimes apply universally, even to users outside those countries. This phenomenon, termed “censorship-by-proxy,” was observed across commercial models from Anthropic, DeepSeek, Google, Meta, OpenAI, and xAI. The study found that across these ten models, the average refusal rate for requests involving political criticism was 34% in countries with restrictive speech laws, compared to 14% in more permissive nations.
Specific instances highlighted this trend. Anthropic’s Claude Sonnet 4, for example, refused all five requests to generate protest flyers criticizing Xi Jinping, Saudi Crown Prince Mohammed bin Salman, and Thailand’s King Vajiralongkorn. Yet, it readily produced flyers critical of President Donald Trump and King Charles III. Google’s Gemini 3 Pro and Meta’s open-weight Llama 4 Maverick exhibited similar patterns, often invoking local laws or deeming criticism of certain leaders as “sensitive or illegal” when refusing requests. These findings underscore a critical concern: the potential for foundational AI models to reflect and inadvertently entrench the restrictive speech norms of repressive governments, a development that could subtly shape global information access and understanding.
