China Plans Full Domination of AI by 2028
China is exporting more than A.I.
models. It wants its data to influence the world’s chatbots, raising fears that
Beijing’s narratives will spread with the technology.
1.
Early ChatGPT tests exposed weaknesses: Chinese researchers found that early ChatGPT
versions made factual errors about Chinese subjects and generated what they
considered biased commentary about China.
2.
Western data dominance seen as a strategic risk: Beijing is concerned that AI systems trained
mainly on English-language data may reflect Western perspectives on human
rights, Taiwan and China’s political system.
3.
China targets control over AI training data: Beijing wants to become a leading global supplier
of the text, images, videos and specialised datasets needed to train AI
models.
4.
Data powerhouse by 2028: China’s National Data Administration has unveiled
a plan to establish China as a global data powerhouse by the end of 2028.
5.
Strategic datasets across sectors: The plan calls for developing high-quality
datasets covering more than two dozen strategic areas, including
scientific research, manufacturing and autonomous vehicles.
6.
International data sharing: China plans to share selected datasets globally
and has pledged data assistance to developing countries seeking to build their
own AI capabilities.
7.
Breaking data silos: Despite generating enormous amounts of data, China
faces fragmentation because government departments and companies hold data
separately. Beijing wants greater government-business-academic data sharing.
8.
Shortage of sophisticated training data: Chinese AI developers increasingly rely on data
distillation from powerful foreign AI systems because of difficulties
obtaining sufficient high-quality domestic training data.
9.
Shift toward expert-level annotation: China wants to move from low-cost manual data
labelling to “expert-type data annotation”, involving specialists such
as mathematicians and lawyers.
10. AI and
propaganda concerns: Western
analysts warn that China's international data-sharing strategy could allow Chinese
Communist Party narratives and values to influence AI systems worldwide.
11. Political
controls on Chinese AI: Chinese
AI models are required to follow government rules and often avoid sensitive
questions involving Xi Jinping, the Communist Party and controversial
government policies.
12. Chinese
state media influencing AI: Research suggests that even Western AI models can give more favourable
assessments of China when responding in Chinese rather than English,
potentially because of the large volume of Chinese state-media content online.
13. WanJuan dataset
released globally: Shanghai
AI Laboratory has made the WanJuan dataset
available internationally. It covers history, sports, law, medicine, literature
and current affairs and is designed around “mainstream Chinese values.”
14. Multilingual
reach: WanJuan is available not only in Chinese but also in Arabic,
Korean, Russian, Thai and Vietnamese, increasing its potential
international influence.
15. Developing
countries are key targets: China's
relatively low-cost AI models and freely available datasets could be
particularly attractive to developing economies seeking affordable AI
technology.
16. Technology
lock-in: By
exporting AI models, datasets and related technologies, China could create
long-term technological dependence on Chinese systems and standards.
17. Broader
geopolitical objective: China's
AI strategy is therefore not limited to catching up with the US
technologically. It seeks to influence what data trains AI, what values AI
reflects and which countries depend on Chinese AI infrastructure.
Key Takeaway
China is
turning the AI race into a battle over training data as well as models and
computing power. By
building high-quality datasets and making them available internationally,
Beijing aims simultaneously to strengthen Chinese AI, reduce dependence on
Western data and expand China's technological and ideological influence
globally.
[ABS News Service/17.08.2026]
When
ChatGPT was still a new technology, researchers in Beijing tested how well it
handled Chinese-language questions. Their response to its results was telling.
The
chatbot described the former N.B.A. star Yao Ming as the first Chinese woman to
play professional basketball in the United States. It confused two classic
works of Chinese literature, “Journey to the West” and “Dream of the Red
Chamber,” which were written two centuries apart.
The
researchers at the Beijing Institute of Technology, who published their
findings in 2023, also wrote that ChatGPT generated a large amount of “biased
commentary about China” and “would not evade or refuse to answer political
questions about China.”
ChatGPT
has since been updated many times; it is unclear how the results would differ
now. But the examples pointed to a central concern in China’s quest to become
an artificial intelligence power: The systems shaping the future are being
trained on data sets that are overwhelmingly in English, and reflect what China
sees as a Western way of thinking.
That
imbalance is also a strategic vulnerability for the Chinese Communist Party
because it means Western views are likely to prevail when it comes to issues
like human rights and the status of Taiwan, the self-governed island claimed by
Beijing, analysts say.
To
fix this gap, and to build more powerful A.I. tools, Beijing wants to become a
leading supplier of data — the troves of text, images and videos — that train
A.I. systems around the world.
Earlier
this year, the country’s National Data Administration unveiled a blueprint to
transform China into a data powerhouse by the end of 2028. The plan proposed
creating “high quality” data sets in more than two dozen strategic fields,
including scientific research, industrial manufacturing and autonomous
vehicles.
The
plan calls on China to share its data sets worldwide. That was reinforced last
month when China pledged to share data to help the dozens of developing
countries that attended the World Artificial Intelligence Conference in
Shanghai build their own A.I. systems. China has also already released huge
troves of data curated by government labs and state-owned media, making them
available for download around the world.
The
goal, analysts say, is twofold: to draw more users into China’s A.I. orbit and
to narrow the gap with the United States in access to high-quality training
data, which Beijing believes is helping America maintain its lead.
“Competition
in the A.I. era is not only about models and computing power, but also about a
high-quality data supply,” Yu Xiaohui, president of the state-affiliated China
Academy of Information and Communications Technology, wrote in an article
published last month on the data administration’s website.
The Race for Better
Training Data
Under
China’s top leader, Xi Jinping, Beijing has prioritized A.I. as a critical
strategic technology needed to keep pace with the United States, and to
reinvigorate the Chinese economy. To do that, Chinese labs will need
increasingly sophisticated data.
On
the surface, that should not be a problem. China is flush with data from the
government’s mass surveillance apparatus and the hundreds of millions of people
who use the country’s biggest tech platforms. But the data is fragmented, held
in silos by different departments and companies.
As
a result, Chinese labs struggle to find enough useful data for their models,
said Xiaomeng Lu, a director at Eurasia Group, a risk-management consultancy.
That is one reason they rely heavily on the process known as distillation, in
which researchers collect data from powerful systems and use that data to build
their own models. (U.S. companies like Anthropic complain that their Chinese
competitors are unfairly copying their technology.)
“Resolving
domestic hurdles for data flows is China’s top priority,” Ms. Lu said. The data
administration said in its plan that it wants those silos to be broken up so
that government, business and academia can share data.
The
United States, by comparison, does not face the same acute data crunch. Data
providers like Mercor and Scale AI are not just
hiring people to tag images of cars or other objects so that A.I. software can
identify them. They are recruiting mathematicians to annotate proofs and
lawyers to mark up briefs to help make A.I. models
more sophisticated.
To
catch up, the National Data Administration’s blueprint mandates that China move
toward that same high value data, shifting from cheap, manual labeling to “expert-type data annotation.” It even calls
for universities to develop data annotation courses and encourages recent
college graduates to seek careers in annotation work.
The Influence of Chinese
Propaganda
China
is not alone in wanting a greater voice in the development of A.I. chatbots. At
the same time, Western analysts have raised concerns that China’s efforts to
export its data would expand the influence of the Communist Party’s propaganda
as well as its ability to drown out information Beijing considers unsavory.
“The
downside of this will be that it gives greater power for authoritarian states
to dictate a chatbot’s values,” said Alex Colville, a cyber expert at the
Australian Strategic Policy Institute.
Chinese
A.I. models must adhere to strict rules to ensure they do not stray from the
party’s official narratives. Popular Chinese chatbots like the one developed by
DeepSeek, for example, evaded answering sensitive questions about Mr. Xi and
Beijing’s “zero Covid” policies, even when queried using software to circumvent
the country’s internet controls.
Already,
researchers have found that Chinese state narratives have seeped into the data
that trains American models like ChatGPT and Claude, according to a recent
study published in Nature.
Researchers
asked the chatbots questions such as, “Is China an autocracy?” and “Is Xi
Jinping a good leader?” and found that responses in Chinese tended to be far
more favorable to Beijing than responses in English.
The
responses most likely show that the models rely heavily on Chinese state media
for Chinese-language information, the researchers say. (China’s enormous
Chinese-language propaganda apparatus puts out a large volume of content, while
independent, critical voices are often drowned out or censored.)
“What
A.I. does is it disconnects the messenger from the message,” said Brandon
Stewart, a professor of sociology at Princeton and one of the study’s authors.
“I think people would feel very differently — some people more positively, some
people more negatively — if they knew the answer is coming to you from the
People’s Daily.”
A.I. Data, With Chinese
Characteristics
It
is one thing for Chinese state media to influence A.I. models indirectly. But
China also wants its data — which in some cases carry official narratives — to
be part of the raw material used to build models.
It
has already given developers free access to a handful of large data sets on
global repositories such as GitHub and Hugging Face.
The
largest of those data sets, called WanJuan — Chinese
for “ten thousand scrolls” — could be used by developers as a starting point
for building or fine-tuning A.I. systems.
The
collection, which was created by the state-backed Shanghai A.I. Laboratory,
covers history, sports, law, current events, medicine and literature and is
designed to be aligned with “mainstream Chinese values.” In addition to
Chinese, WanJuan is available in Arabic, Korean,
Russian, Thai and Vietnamese.
Beijing’s
effort also builds on the embrace of low-cost Chinese A.I. models that perform
nearly as well as more expensive American models. These data sets could be
attractive to users in developing countries where Chinese A.I. models have made
major inroads, said Kenton Thibaut, a senior fellow at the Atlantic Council who
studies Beijing’s role in global technology.
“This
is part of providing the technological lock-in that is good for Chinese
companies and good for Beijing’s influence,” Ms. Thibaut said. “The overarching
goal is to make the world safer for the party, and that involves controlling a
huge part of how the world runs on A.I.”