China Plans Full Domination of AI by 2028

China is exporting more than A.I. models. It wants its data to influence the world’s chatbots, raising fears that Beijing’s narratives will spread with the technology.

1.    Early ChatGPT tests exposed weaknesses: Chinese researchers found that early ChatGPT versions made factual errors about Chinese subjects and generated what they considered biased commentary about China.

2.    Western data dominance seen as a strategic risk: Beijing is concerned that AI systems trained mainly on English-language data may reflect Western perspectives on human rights, Taiwan and China’s political system.

3.    China targets control over AI training data: Beijing wants to become a leading global supplier of the text, images, videos and specialised datasets needed to train AI models.

4.    Data powerhouse by 2028: China’s National Data Administration has unveiled a plan to establish China as a global data powerhouse by the end of 2028.

5.    Strategic datasets across sectors: The plan calls for developing high-quality datasets covering more than two dozen strategic areas, including scientific research, manufacturing and autonomous vehicles.

6.    International data sharing: China plans to share selected datasets globally and has pledged data assistance to developing countries seeking to build their own AI capabilities.

7.    Breaking data silos: Despite generating enormous amounts of data, China faces fragmentation because government departments and companies hold data separately. Beijing wants greater government-business-academic data sharing.

8.    Shortage of sophisticated training data: Chinese AI developers increasingly rely on data distillation from powerful foreign AI systems because of difficulties obtaining sufficient high-quality domestic training data.

9.    Shift toward expert-level annotation: China wants to move from low-cost manual data labelling to “expert-type data annotation”, involving specialists such as mathematicians and lawyers.

10.  AI and propaganda concerns: Western analysts warn that China's international data-sharing strategy could allow Chinese Communist Party narratives and values to influence AI systems worldwide.

11.  Political controls on Chinese AI: Chinese AI models are required to follow government rules and often avoid sensitive questions involving Xi Jinping, the Communist Party and controversial government policies.

12.  Chinese state media influencing AI: Research suggests that even Western AI models can give more favourable assessments of China when responding in Chinese rather than English, potentially because of the large volume of Chinese state-media content online.

13.  WanJuan dataset released globally: Shanghai AI Laboratory has made the WanJuan dataset available internationally. It covers history, sports, law, medicine, literature and current affairs and is designed around “mainstream Chinese values.”

14.  Multilingual reach: WanJuan is available not only in Chinese but also in Arabic, Korean, Russian, Thai and Vietnamese, increasing its potential international influence.

15.  Developing countries are key targets: China's relatively low-cost AI models and freely available datasets could be particularly attractive to developing economies seeking affordable AI technology.

16.  Technology lock-in: By exporting AI models, datasets and related technologies, China could create long-term technological dependence on Chinese systems and standards.

17.  Broader geopolitical objective: China's AI strategy is therefore not limited to catching up with the US technologically. It seeks to influence what data trains AI, what values AI reflects and which countries depend on Chinese AI infrastructure.

Key Takeaway

China is turning the AI race into a battle over training data as well as models and computing power. By building high-quality datasets and making them available internationally, Beijing aims simultaneously to strengthen Chinese AI, reduce dependence on Western data and expand China's technological and ideological influence globally.

 

[ABS News Service/17.08.2026]

When ChatGPT was still a new technology, researchers in Beijing tested how well it handled Chinese-language questions. Their response to its results was telling.

The chatbot described the former N.B.A. star Yao Ming as the first Chinese woman to play professional basketball in the United States. It confused two classic works of Chinese literature, “Journey to the West” and “Dream of the Red Chamber,” which were written two centuries apart.

The researchers at the Beijing Institute of Technology, who published their findings in 2023, also wrote that ChatGPT generated a large amount of “biased commentary about China” and “would not evade or refuse to answer political questions about China.”

ChatGPT has since been updated many times; it is unclear how the results would differ now. But the examples pointed to a central concern in China’s quest to become an artificial intelligence power: The systems shaping the future are being trained on data sets that are overwhelmingly in English, and reflect what China sees as a Western way of thinking.

That imbalance is also a strategic vulnerability for the Chinese Communist Party because it means Western views are likely to prevail when it comes to issues like human rights and the status of Taiwan, the self-governed island claimed by Beijing, analysts say.

To fix this gap, and to build more powerful A.I. tools, Beijing wants to become a leading supplier of data — the troves of text, images and videos — that train A.I. systems around the world.

Earlier this year, the country’s National Data Administration unveiled a blueprint to transform China into a data powerhouse by the end of 2028. The plan proposed creating “high quality” data sets in more than two dozen strategic fields, including scientific research, industrial manufacturing and autonomous vehicles.

The plan calls on China to share its data sets worldwide. That was reinforced last month when China pledged to share data to help the dozens of developing countries that attended the World Artificial Intelligence Conference in Shanghai build their own A.I. systems. China has also already released huge troves of data curated by government labs and state-owned media, making them available for download around the world.

The goal, analysts say, is twofold: to draw more users into China’s A.I. orbit and to narrow the gap with the United States in access to high-quality training data, which Beijing believes is helping America maintain its lead.

“Competition in the A.I. era is not only about models and computing power, but also about a high-quality data supply,” Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, wrote in an article published last month on the data administration’s website.

The Race for Better Training Data

Under China’s top leader, Xi Jinping, Beijing has prioritized A.I. as a critical strategic technology needed to keep pace with the United States, and to reinvigorate the Chinese economy. To do that, Chinese labs will need increasingly sophisticated data.

On the surface, that should not be a problem. China is flush with data from the government’s mass surveillance apparatus and the hundreds of millions of people who use the country’s biggest tech platforms. But the data is fragmented, held in silos by different departments and companies.

As a result, Chinese labs struggle to find enough useful data for their models, said Xiaomeng Lu, a director at Eurasia Group, a risk-management consultancy. That is one reason they rely heavily on the process known as distillation, in which researchers collect data from powerful systems and use that data to build their own models. (U.S. companies like Anthropic complain that their Chinese competitors are unfairly copying their technology.)

“Resolving domestic hurdles for data flows is China’s top priority,” Ms. Lu said. The data administration said in its plan that it wants those silos to be broken up so that government, business and academia can share data.

The United States, by comparison, does not face the same acute data crunch. Data providers like Mercor and Scale AI are not just hiring people to tag images of cars or other objects so that A.I. software can identify them. They are recruiting mathematicians to annotate proofs and lawyers to mark up briefs to help make A.I. models more sophisticated.

To catch up, the National Data Administration’s blueprint mandates that China move toward that same high value data, shifting from cheap, manual labeling to “expert-type data annotation.” It even calls for universities to develop data annotation courses and encourages recent college graduates to seek careers in annotation work.

The Influence of Chinese Propaganda

China is not alone in wanting a greater voice in the development of A.I. chatbots. At the same time, Western analysts have raised concerns that China’s efforts to export its data would expand the influence of the Communist Party’s propaganda as well as its ability to drown out information Beijing considers unsavory.

“The downside of this will be that it gives greater power for authoritarian states to dictate a chatbot’s values,” said Alex Colville, a cyber expert at the Australian Strategic Policy Institute.

Chinese A.I. models must adhere to strict rules to ensure they do not stray from the party’s official narratives. Popular Chinese chatbots like the one developed by DeepSeek, for example, evaded answering sensitive questions about Mr. Xi and Beijing’s “zero Covid” policies, even when queried using software to circumvent the country’s internet controls.

Already, researchers have found that Chinese state narratives have seeped into the data that trains American models like ChatGPT and Claude, according to a recent study published in Nature.

Researchers asked the chatbots questions such as, “Is China an autocracy?” and “Is Xi Jinping a good leader?” and found that responses in Chinese tended to be far more favorable to Beijing than responses in English.

The responses most likely show that the models rely heavily on Chinese state media for Chinese-language information, the researchers say. (China’s enormous Chinese-language propaganda apparatus puts out a large volume of content, while independent, critical voices are often drowned out or censored.)

“What A.I. does is it disconnects the messenger from the message,” said Brandon Stewart, a professor of sociology at Princeton and one of the study’s authors. “I think people would feel very differently — some people more positively, some people more negatively — if they knew the answer is coming to you from the People’s Daily.”

A.I. Data, With Chinese Characteristics

It is one thing for Chinese state media to influence A.I. models indirectly. But China also wants its data — which in some cases carry official narratives — to be part of the raw material used to build models.

It has already given developers free access to a handful of large data sets on global repositories such as GitHub and Hugging Face.

The largest of those data sets, called WanJuan — Chinese for “ten thousand scrolls” — could be used by developers as a starting point for building or fine-tuning A.I. systems.

The collection, which was created by the state-backed Shanghai A.I. Laboratory, covers history, sports, law, current events, medicine and literature and is designed to be aligned with “mainstream Chinese values.” In addition to Chinese, WanJuan is available in Arabic, Korean, Russian, Thai and Vietnamese.

Beijing’s effort also builds on the embrace of low-cost Chinese A.I. models that perform nearly as well as more expensive American models. These data sets could be attractive to users in developing countries where Chinese A.I. models have made major inroads, said Kenton Thibaut, a senior fellow at the Atlantic Council who studies Beijing’s role in global technology.

“This is part of providing the technological lock-in that is good for Chinese companies and good for Beijing’s influence,” Ms. Thibaut said. “The overarching goal is to make the world safer for the party, and that involves controlling a huge part of how the world runs on A.I.”