CHECKED
首个中文COVID-19假新闻多模态数据集,包含2104条经核查的微博,用于研究疫情虚假信息的传播与内容分析。
基本信息
资源简介
CHECKED是首个中文COVID-19假新闻数据集,由雪城大学数据实验室创建,包含2104条经过验证的微博(2019年12月至2020年8月),分为真实和虚假两类。数据涵盖文本、视觉、时间、网络等多种模态,以及转发、评论、点赞数,用于研究COVID-19假新闻的传播模式和内容分析,提升公众对疫情信息的辨识能力。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/cyang03/CHECKED.git
curl -L -o repo.zip https://github.com/cyang03/CHECKED/archive/refs/heads/master.zip
unzip repo.zip
源站 README 摘录(使用方式)
CHECKED
The first Chinese COVID-19 fake news dataset based on the Weibo platform. Please check out our paper here.
Notice
We care about users’’ privacy and made (will keep making) efforts to protecting it.
- For microblogs: We released the hashed
idinstead of the originalidof microblogs. - For users: We did not make the
user_namepublic, which enables to identify Weibo users. In addition, we released the hasheduser_idinstead of the originaluser_id. - Please use the CHECKED data only for academic research.
Update:
We corrected a small amount of dates displayed in the comment sections of two microblogs, which occured due to unknown errors during the automatic information extraction process. We accomplished this by checking the dataset manually. We appreciate the user for pointing this issue out. If similar problems are to be found by future users, please contact us as we will fix them immediately. (August 22, 2021)
We addanalysisas a new keyword for each microblog labeled as fake.analysiscontains the expert analysis and justification, which details the news falseness. Please check out the dataset folder for details. We also provide benchmark results of FastText, TextCNN, TextRNN, Att-TextRNN, and Transformer using the CHECKED data in predicting fake news. Please check out the baseline folder for details. (June 9, 2021)
Overiew
This repository includes three folders, named as (1) dataset, (2) code, (3) baseline, respectively.
- dataset: This folder contains the CHECKED data in both
jsonandcsvformat and the list of keywords used to determine whether a microblog is relevant to COVID-19 or not. Specifically,- fake_news: This folder includes 344 fake microblogs (in
jsonformat). - real_news: This folder includes 1760 real microblogs (in
jsonformat). - *.csv: These
csvfiles are converted fromjsonfiles in the fake_news and real_news folder. - keyword_list.txt: This file includes all the keywords that we use t
- fake_news: This folder includes 344 fake microblogs (in
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。




