A massive data warehouse containing information on about one billion Chinese residents could be one of the biggest breaches of personal information in history.
Parts of the leaked data appeared last week on a prominent cybercrime forum from someone selling the cache for 10 bitcoins, or about $200,000, and were said to have been pulled from a Shanghai police database stored on Alibaba’s cloud.
Although details of the breach remain scarce, parts of the data have been verified as authentic, suggesting that at least some of the data is real. The origin of the data and how it ended up in the hands of an underground seller, whose motives are unknown, is still unclear.
News of the alleged breach went largely unnoticed in mainland China, where restrictions on speech and expression are tightly controlled and internet access is censored and severely restricted.
The breach, if authentic, raises questions about the sheer scale of China’s surveillance state, the largest and most expansive in the world, and Beijing’s ability to keep that data safe.
Here’s what we’ve learned so far.
How did the data leak?
In a since-deleted cybercrime forum post, the seller claimed to have downloaded the data from a cloud storage server hosted by Alibaba, the cloud computing division of the Chinese e-commerce giant. When contacted by TechCrunch on Monday, Alibaba said it was looking into the allegations.
Exactly how the data was leaked is hazy, but experts say the database may have been misconfigured and exposed due to human error since April 2021 before it was discovered. This appears to rule out the claim that the database credentials were inadvertently published as part of a technical blog post on a Chinese developer site in 2020 and later used to extract the billion records from the police database, as it did not passwords were required to access it.
Bob Diachenko, a Ukrainian security researcher, told TechCrunch that his own monitoring records showed that the database was also exposed through a Kibana dashboard, a web-based software used to visualize and search massive Elasticsearch databases, in the end of April. If the database did not require a password, as was believed, anyone could access the data if they knew its web address.
Security researchers often scan the Internet for inadvertently exposed databases or other sensitive data, often to collect bounties offered by the companies they help secure. But threat actors also perform the same scans, often with the goal of copying data from an exposed database, deleting it, and offering to return the data in return for payment of a ransom — an increasingly common tactic used by container-diving criminals in recent years. Dyachenko said this is what happened in this case; a malicious actor found, attacked and deleted the exposed database and left behind a ransom note demanding 10 bitcoins for its return.
“My hypothesis here is that the ransom note didn’t work and the threat decided to get money somewhere else. Or another malicious player came across the data and decided to put it up for sale,” Dyachenko said.
Little is known about the seller or for what reason the data was dumped online. It’s not unusual to see large amounts of personal data for sale on cybercrime forums and the dark web, but rarely for data this sensitive or in this amount.
What does the data look like?
TechCrunch reviewed a larger sample of the data uploaded by the seller, which contained three files totaling about 500 megabytes, each containing 250,000 individual records.
The data itself is formatted in JSON, a standard file format for Elasticsearch databases, making it easy to read and analyze. The format of the database suggests that it is meticulously maintained and downloaded, rather than created by purely aggregating information from multiple data sources, a common technique used by information vendors and data brokers. However, some data may be derived from external sources, such as food delivery orders.
What also makes the data likely to be genuine is the sheer size of the data and that level of detail would be difficult – though not impossible – to falsify.
TechCrunch translated the police records, which were written in Chinese, and redacted the personal information.
The files appear to contain detailed police reports from 1995 to 2019, including names, addresses, phone numbers, social security numbers, gender, and the reason police were called. The records seen by TechCrunch include detailed coordinates of where incidents occurred or police reports were made — and the names of the informants who made the reports — that match the exact addresses also listed in each record, as well as the race and ethnicity of the individuals. (The Chinese government has imprisoned more than a million of its citizens, mostly from Muslim minority ethnic groups, including Uyghurs and Kazakhs, in what the Biden administration has called “genocide.”)
The records contain complaints and criminal charges, from serious crimes involving violence to the relatively trivial, such as detailed reports of credit card fraud, Internet scams and gambling, which is illegal in China. Several records seen by TechCrunch show police reports prosecuting the use of VPNs, or virtual private networks, used to access sites blocked by China’s censorship system and as such banned in China. One recording shows a Shanghai resident accused of using a VPN to post remarks critical of the government on Twitter, which is banned in China. It is not known what happened to the person afterwards.
The data also contained full web addresses of photos stored on the same server, none of which were available at the time of writing, but the associated data often shows what was uploaded, such as a person’s residence document or their passport when leaving the country. These web addresses are formatted in a way that is consistent with how Alibaba’s cloud service stores files.
Many of the records we reviewed appeared to contain information about children based on their birth dates and ages listed in the records.
Without (unlikely) confirmation from the Chinese government, it’s hard to know for sure if the seller’s claims are true and the data was obtained from the Shanghai Police Department as claimed. The Wall Street Journal, The New York Times and CNN have verified portions of the data by calling individuals whose information was found in the database, lending weight to its authenticity.
What is the impact?
This alleged breach, if proven to be legitimate, could be very damaging for Beijing and raises questions about the government’s cybersecurity measures and the impact the breach will have on people.
It comes at a time when China is strengthening privacy protections. Last September, China passed the Personal Information Protection Law, the first comprehensive privacy and data protection legislation, widely seen as the Chinese equivalent of Europe’s GDPR privacy rules. The law limits how companies can collect personal data and is expected to have a huge effect on the advertising business of the country’s biggest tech giants, but allows major exemptions for government agencies and departments that make up China’s vast surveillance capabilities.
Beijing is already reportedly censoring news of the alleged breach, with Chinese messaging apps WeChat and Weibo blocking messages and mentions such as “data leak” and “database breach.” The Chinese government has yet to comment on the breach.
This is not the first security breach involving a vast trove of data on Chinese residents that has been left exposed to the wider internet without a password. In 2019, TechCrunch reported that a smart city installation in China was dispersing the contents of a facial recognition database to nearby residents.
You can reach this reporter on Signal and WhatsApp at +1 646-755-8849 or zack.whittaker@techcrunch.com via email.
Add Comment