Coswara呼吸声音数据集
Coswara 是由 IISc 发起的呼吸声音数据集,包含咳嗽、快慢呼吸、元音和语音等多类音频及健康元数据,适合哮喘、咳嗽和呼吸疾病相关声音分类与多模态筛查。
基本信息
资源简介
Coswara 收集咳嗽、呼吸和语音音频及症状元数据,可用于呼吸疾病远程筛查。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/iiscleap/Coswara-Data.git
curl -L -o repo.zip https://github.com/iiscleap/Coswara-Data/archive/refs/heads/master.zip
unzip repo.zip
源站 README 摘录(使用方式)
Coswara-Data
<strong>UPDATE:</strong> The current full version of Coswara data is now published with open access in <strong>Nature Scientific Data</strong>, 2023. read
Project Coswara by Indian Institute of Science (IISc) Bangalore is an attempt to build a diagnostic tool for COVID-19 detection using the audio recordings such as breathing, cough and speech sounds of an individual. Currently, the project is in the data collection stage through crowdsourcing. To contribute your audio samples, please go to Project Coswara(https://coswara.iisc.ac.in/). The exercise takes 5-7 minutes.
<strong>What am I looking at?</strong>
This github repository contains the raw audio data collected through https://coswara.iisc.ac.in/ . Every participant contributes nine sound samples. You can read the paper: Coswara - A Database of Breathing, Cough, and Voice Sounds for COVID-19 Diagnosis to know more about the dataset. Note that the dataset size has increased since this paper came out. We also maintain a (less frequently updated) blog here.
<strong>What is the structure of the repository?</strong>
Each folder contains metadata and audio recordings corresponding to contributors. The folder is compressed. To download and extract the data, you can run the script extract_data.py
<strong>What are the different sound samples?</strong>
Sound samples collected include breathing sounds (fast and slow), cough sounds (deep and shallow), phonation of sustained vowels (/a/ as in made, /i/,/o/), and counting numbers at slow and fast pace. Metadata information collected includes the participant’'s age, gender, location (country, state/ province), current health status (healthy/ exposed/ positive/recovered) and the presence of comorbidities (pre-existing medical conditions).
<strong>Can I see the metadata before downloading whole repository?</strong>
Yes. The file combined_data.csv contains a summary of metadata. The file csv_labels_legend.json contains information about the columns present in combined_data.csv.
<strong>Is there any audio quality check?</strong>
Yes. The audio files are manually listened and labeled as one of the three categories: 2(excellent), 1(good), 0(bad). The labels are presen
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 运动诱发性哮喘 的高质量、多模态真实临床数据定制解决方案。




