COVID-TAD

COVID-TAD是一个大规模COVID-19错误信息数据集,覆盖25个月,包含多个时间快照的社交媒体帖子,用于检测和分析错误信息。

佐治亚理工学院计算机科学学院佐治亚理工学院计算机科学学院
arXiv
2022-11-22 更新
浏览 7
文本COVID-19错误信息检测

基本信息

模态
文本
创建/更新时间
2022-11-22

资源简介

COVID-TAD是一个大规模COVID-19错误信息数据集,覆盖25个月时间跨度,包含多个时间快照的社交媒体帖子,用于检测和分析COVID-19相关的错误信息。

原始链接

http://arxiv.org/abs/2211.12508v1

访问原始数据

官方服务

如需原始数据获取支持或标注服务,请联系我们。

帮我联系

下载信息

注册下载

Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。

暂未开放

公开下载

Tips: 该数据集属于公开下载,应该可以免费公开下载。

免登录

有偿下载

Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。

提供高速下载与技术交付服务(收技术服务费,非数据销售)

暂未开放

千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。

使用方式

数据集说明

COVID-TAD 对应论文数据集(arXiv 预印本)。

数据获取指引

  1. 打开论文页面获取作者与项目信息:https://arxiv.org/abs/2211.12508v1
  2. 论文 Data Availability / Code Availability 章节标注了数据实际托管位置;
  3. 获取到实际数据链接后,按对应平台标准方式下载。

论文摘要:Abstract:Recent advances in text classification and knowledge capture in language models have relied on availability of large-scale text datasets. However, language models are trained on static snapshots of knowledge and are limited when that knowledge evolves. This is especially critical for misinformation detection, where new types of misinformation continuously appear, replacing old campaigns. We propose time-aware misinformation datasets to capture time-critical phenomena. In this paper, we first present evidence of evolving misinformation and show that incorporating even simple time-awareness significantly improves classifier accuracy. Second, we present COVID-TAD, a large-scale COVID-19 misinformation da-taset spanning 25 months. It is the first large-scale misinformation dataset that contains multiple snapshots of a datastream and is orders of magnitude bigger than related misinformation datasets. We describe the collection and labeling pro-cess, as well as preliminary experiments.

论文页面:https://arxiv.org/abs/2211.12508v1

精度瓶颈?数据缺失?

当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。

获取专属数据定制方案