CORD19STS

COVID-19语义文本相似性数据集,包含13,710个标注句子对,用于支持对话诊断和信息检索。

美国南加州大学信息科学研究所美国南加州大学信息科学研究所
arXiv
2020-11-03 更新
浏览 3
文本COVID-19语义文本相似性

基本信息

模态
文本
创建/更新时间
2020-11-03

资源简介

该数据集是专为COVID-19设计的语义文本相似性数据集,由美国南加州大学信息科学研究所创建,包含13,710个标注的句子对,句子对来自COVID-19开放研究数据集(CORD-19)。数据模态为文本,任务为语义文本相似性,用于支持对话医疗诊断系统和信息检索等自然语言处理应用。

原始链接

http://arxiv.org/abs/2007.02461v2

访问原始数据

官方服务

如需原始数据获取支持或标注服务,请联系我们。

帮我联系

下载信息

注册下载

Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。

暂未开放

公开下载

Tips: 该数据集属于公开下载,应该可以免费公开下载。

免登录

有偿下载

Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。

提供高速下载与技术交付服务(收技术服务费,非数据销售)

暂未开放

千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。

使用方式

数据集说明

CORD19STS 对应论文数据集(arXiv 预印本)。

数据获取指引

  1. 打开论文页面获取作者与项目信息:https://arxiv.org/abs/2007.02461v2
  2. 论文 Data Availability / Code Availability 章节标注了数据实际托管位置;
  3. 获取到实际数据链接后,按对应平台标准方式下载。

论文摘要:Abstract:In order to combat the COVID-19 pandemic, society can benefit from various natural language processing applications, such as dialog medical diagnosis systems and information retrieval engines calibrated specifically for COVID-19. These applications rely on the ability to measure semantic textual similarity (STS), making STS a fundamental task that can benefit several downstream applications. However, existing STS datasets and models fail to translate their performance to a domain-specific environment such as COVID-19. To overcome this gap, we introduce CORD19STS dataset which includes 13,710 annotated sentence pairs collected from COVID-19 open research dataset (CORD-19) challenge. To be specific, we generated one million sentence pairs using different sampling strategies. We then used a finetuned BERT-like language model, which we call Sen-SCI-CORD19-BERT, to calculate the similarity scores between sentence pairs to provide a balanced dataset with respect to the different semantic similarity levels, which gives us a total of 32K sentence pairs. Each sentence pair was annotated by five Amazon Mechanical Turk (AMT) crowd workers, where the labels represent different semantic similarity levels between the sentence pairs (i.e. related, somewhat-related, and not-related). After employing a rigorous qualification tasks to verify collected annotations, our final CORD19STS dataset includes 13,710 sentence pairs.

论文页面:https://arxiv.org/abs/2007.02461v2

精度瓶颈?数据缺失?

当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。

获取专属数据定制方案