‘Multi-part case study on China’s media’ finds that AI models can’t hallucinate away Chinese censorship
Regulators and researchers spent the past year warning that Chinese AI models like DeepSeek come with Beijing’s censorship baked in , but a growing body of research suggests American AI models may not be entirely immune from it.
A peer-reviewed study published in Nature found evidence that Chinese state-controlled media makes its way into AI training data and can influence how models answer questions about China.
Meta’s independent Oversight Board found models from Anthropic, OpenAI, Google and Meta were more than twice as likely to refuse requests to criticize governments in countries that restrict political speech than those in freer countries—raising questions about whether the rules of authoritarian information environments are migrating into AI products used around the world.
“This kind of influence is concerning because—like covert information operations—it severs information and opinion from their source, effectively laundering government-manipulated content into ostensibly objective text,” the Nature researchers wrote.
The evidence taken together suggests the effects of China’s tightly controlled information system may not stop at Chinese-built AI, though neither shows that Beijing deliberately manipulated OpenAI, Anthropic, Google or Meta.
It does point to a deeper problem for an industry that markets its models as politically neutral.
Anthropic, for example, has touted efforts to make Claude politically “even-handed,” reporting a 94% score on its own evaluation.
How Chinese propaganda seeps into training data The Nature paper researchers identified more than three million Chinese-language documents in the open-source training dataset CulturaX–used to train and improve LLMs— and built a “multi-part case study on China’s media” that particularly focused on political subjects.
They also found signs that the models had encountered the material during training: Claude Sonnet, Claude Opus, GPT-3.5 Instruct, GPT-4 and GPT-4o could reproduce distinctive phrases from Chinese state-coordinated media at rates ranging from 3% to nearly 10%.
The researchers then tested what that kind of material could do to a model.
After further training Meta’s open-weight Llama 2 13B—picked because it had very little to zero Chinese state media in its training data—on just 6,400 Chinese state-scripted news examples, the model produced a more Beijing-friendly answer than the baseline model nearly 80% of the time.
At higher levels of additional training, the difference became stark.
After 64,000 state-scripted examples, the retrained model was asked whether China is an autocracy.
The baseline model said it was.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.