COVID-19相关推文数据集
包含3049条COVID-19相关推文,标记为真实或虚假,用于分析假新闻的语言特征。
基本信息
资源简介
该数据集包含3049条经过清洗和筛选的COVID-19相关推文,其中2161条标记为真实,888条标记为虚假。数据来源于Twitter官方账号及事实核查网站(如PolitiFact、Poynter、Snopes),由研究助理手动标注。主要用于分析假新闻与真实新闻的语义特征,以通过语言分析提高社交媒体信息的真实性和可信度。数据模态为文本,任务为真假新闻二分类。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集说明
COVID-19相关推文数据集 对应论文数据集(arXiv 预印本)。
数据获取指引
- 打开论文页面获取作者与项目信息:https://arxiv.org/abs/2310.04237v1
- 论文 Data Availability / Code Availability 章节标注了数据实际托管位置;
- 获取到实际数据链接后,按对应平台标准方式下载。
论文摘要:Abstract:This study investigates the linguistic traits of fake news and real news. There are two parts to this study: text data and speech data. The text data for this study consisted of 6420 COVID-19 related tweets re-filtered from Patwa et al. (2021). After cleaning, the dataset contained 3049 tweets, with 2161 labeled as 'real' and 888 as 'fake'. The speech data for this study was collected from TikTok, focusing on COVID-19 related videos. Research assistants fact-checked each video's content using credible sources and labeled them as 'Real', 'Fake', or 'Questionable', resulting in a dataset of 91 real entries and 109 fake entries from 200 TikTok videos with a total word count of 53,710 words. The data was analysed using the Linguistic Inquiry and Word Count (LIWC) software to detect patterns in linguistic data. The results indicate a set of linguistic features that distinguish fake news from real news in both written and speech data. This offers valuable insights into the role of language in shaping trust, social media interactions, and the propagation of fake news.
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。




