COVYT语音数据集
COVYT数据集包含来自YouTube和TikTok的65名相同说话者感染与未感染COVID-19的语音数据,用于非侵入性COVID-19检测研究。
基本信息
资源简介
COVYT数据集包含来自YouTube和TikTok的8小时以上语音数据,涵盖65名感染与未感染COVID-19的相同说话者,旨在通过语音分析非侵入性地检测COVID-19,适用于机器学习模型训练和个性化COVID-19检测研究。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集说明
COVYT语音数据集 对应论文数据集(arXiv 预印本)。
数据获取指引
- 打开论文页面获取作者与项目信息:https://arxiv.org/abs/2206.11045v1
- 论文 Data Availability / Code Availability 章节标注了数据实际托管位置;
- 获取到实际数据链接后,按对应平台标准方式下载。
论文摘要:Abstract:More than two years after its outbreak, the COVID-19 pandemic continues to plague medical systems around the world, putting a strain on scarce resources, and claiming human lives. From the very beginning, various AI-based COVID-19 detection and monitoring tools have been pursued in an attempt to stem the tide of infections through timely diagnosis. In particular, computer audition has been suggested as a non-invasive, cost-efficient, and eco-friendly alternative for detecting COVID-19 infections through vocal sounds. However, like all AI methods, also computer audition is heavily dependent on the quantity and quality of available data, and large-scale COVID-19 sound datasets are difficult to acquire amongst other reasons due to the sensitive nature of such data. To that end, we introduce the COVYT dataset a novel COVID-19 dataset collected from public sources containing more than 8 hours of speech from 65 speakers. As compared to other existing COVID-19 sound datasets, the unique feature of the COVYT dataset is that it comprises both COVID-19 positive and negative samples from all 65 speakers. We analyse the acoustic manifestation of COVID-19 on the basis of these perfectly speaker characteristic balanced `in-the-wild' data using interpretable audio descriptors, and investigate several classification scenarios that shed light into proper partitioning strategies for a fair speech-based COVID-19 detection.
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。




