CHECKED

首个中文COVID-19假新闻多模态数据集,包含2104条经核查的微博,用于研究疫情虚假信息的传播与内容分析。

数据实验室,电气工程与计算机科学系,雪城大学数据实验室,电气工程与计算机科学系,雪城大学
arXiv
2021-06-13 更新
浏览 3
多模态COVID-19假新闻检测

基本信息

模态
多模态
创建/更新时间
2021-06-13

资源简介

CHECKED是首个中文COVID-19假新闻数据集,由雪城大学数据实验室创建,包含2104条经过验证的微博(2019年12月至2020年8月),分为真实和虚假两类。数据涵盖文本、视觉、时间、网络等多种模态,以及转发、评论、点赞数,用于研究COVID-19假新闻的传播模式和内容分析,提升公众对疫情信息的辨识能力。

原始链接

https://github.com/cyang03/CHECKED

arXiv 论文 →
访问原始数据

官方服务

如需原始数据获取支持或标注服务,请联系我们。

帮我联系

下载信息

注册下载

Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。

暂未开放

公开下载

Tips: 该数据集属于公开下载,应该可以免费公开下载。

免登录

有偿下载

Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。

提供高速下载与技术交付服务(收技术服务费,非数据销售)

暂未开放

千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。

使用方式

数据集获取

git clone https://github.com/cyang03/CHECKED.git

curl -L -o repo.zip https://github.com/cyang03/CHECKED/archive/refs/heads/master.zip
unzip repo.zip

源站 README 摘录(使用方式)

CHECKED

The first Chinese COVID-19 fake news dataset based on the Weibo platform. Please check out our paper here.

Notice

We care about users’’ privacy and made (will keep making) efforts to protecting it.

  • For microblogs: We released the hashed id instead of the original id of microblogs.
  • For users: We did not make the user_name public, which enables to identify Weibo users. In addition, we released the hashed user_id instead of the original user_id.
  • Please use the CHECKED data only for academic research.
    Update:
    We corrected a small amount of dates displayed in the comment sections of two microblogs, which occured due to unknown errors during the automatic information extraction process. We accomplished this by checking the dataset manually. We appreciate the user for pointing this issue out. If similar problems are to be found by future users, please contact us as we will fix them immediately. (August 22, 2021)
    We add analysis as a new keyword for each microblog labeled as fake. analysis contains the expert analysis and justification, which details the news falseness. Please check out the dataset folder for details. We also provide benchmark results of FastText, TextCNN, TextRNN, Att-TextRNN, and Transformer using the CHECKED data in predicting fake news. Please check out the baseline folder for details. (June 9, 2021)

Overiew

This repository includes three folders, named as (1) dataset, (2) code, (3) baseline, respectively.

  • dataset: This folder contains the CHECKED data in both json and csv format and the list of keywords used to determine whether a microblog is relevant to COVID-19 or not. Specifically,
    • fake_news: This folder includes 344 fake microblogs (in json format).
    • real_news: This folder includes 1760 real microblogs (in json format).
    • *.csv: These csv files are converted from json files in the fake_news and real_news folder.
    • keyword_list.txt: This file includes all the keywords that we use t

数据加载示例(表格/文本类)

import pandas as pd, glob, os

files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
       + glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
       + glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))

完整仓库:github.com/cyang03/CHECKED

精度瓶颈?数据缺失?

当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。

获取专属数据定制方案