China wants to shape what the world’s AI knows

China wants to shape what the world’s AI knows

The New York Times reports:

When ChatGPT was still a new technology, researchers in Beijing tested how well it handled Chinese-language questions. Their response to its results was telling.

The chatbot described the former N.B.A. star Yao Ming as the first Chinese woman to play professional basketball in the United States. It confused two classic works of Chinese literature, “Journey to the West” and “Dream of the Red Chamber,” which were written two centuries apart.

The researchers at the Beijing Institute of Technology, who published their findings in 2023, also wrote that ChatGPT generated a large amount of “biased commentary about China” and “would not evade or refuse to answer political questions about China.”

ChatGPT has since been updated many times; it is unclear how the results would differ now. But the examples pointed to a central concern in China’s quest to become an artificial intelligence power: The systems shaping the future are being trained on data sets that are overwhelmingly in English, and reflect what China sees as a Western way of thinking.

That imbalance is also a strategic vulnerability for the Chinese Communist Party because it means Western views are likely to prevail when it comes to issues like human rights and the status of Taiwan, the self-governed island claimed by Beijing, analysts say.

To fix this gap, and to build more powerful A.I. tools, Beijing wants to become a leading supplier of data — the troves of text, images and videos — that train A.I. systems around the world.

Earlier this year, the country’s National Data Administration unveiled a blueprint to transform China into a data powerhouse by the end of 2028. The plan proposed creating “high quality” data sets in more than two dozen strategic fields, including scientific research, industrial manufacturing and autonomous vehicles.

The plan calls on China to share its data sets worldwide. That was reinforced last month when China pledged to share data to help the dozens of developing countries that attended the World Artificial Intelligence Conference in Shanghai build their own A.I. systems. China has also already released huge troves of data curated by government labs and state-owned media, making them available for download around the world.

The goal, analysts say, is twofold: to draw more users into China’s A.I. orbit and to narrow the gap with the United States in access to high-quality training data, which Beijing believes is helping America maintain its lead.

“Competition in the A.I. ​​era is not only about models and computing power, but also about a high-quality data supply,” Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, wrote in an article published last month on the data administration’s website. [Continue reading…]

Comments are closed.