帕金森病语音数据集
包含195个帕金森病语音特征样本的表格数据集,用于分类检测。
基本信息
资源简介
该数据集来自牛津大学,包含195个实例,其中147个为帕金森病患者,48个为非患者。数据集包含22个语音特征(如频率、音高、振幅/周期等),以及一个标签(1表示帕金森病,0表示非帕金森病),用于帕金森病分类检测任务。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/Aadi-J/Parkinsons-Detection.git
curl -L -o repo.zip https://github.com/Aadi-J/Parkinsons-Detection/archive/refs/heads/main.zip
unzip repo.zip
源站 README 摘录(使用方式)
A Machine Learning Approach for the Diagnosis of Parkinson’'s Disease via Speech Analysis
Introduction
This project, researched in March 2022, aims to provide a machine learning-based approach for accurately diagnosing Parkinson’'s Disease using speech analysis. Due to the deletion of the initial account, the project was reuploaded.
Parkinson’s Disease, the second most prevalent neurodegenerative disorder globally, affects over 10 million people. The current diagnostic methods are only 53% accurate for early diagnosis (within 5 years of symptoms). This project explores a machine learning approach using a dataset from the University of Oxford, focusing on various speech features.
Background
Parkinson’'s Disease
Parkinson’s is characterized by the death of dopamine-containing cells in the substantia nigra, impacting motor and cognitive abilities. Symptoms include frozen facial features, slowness of movement, tremors, and voice impairment.
Performance Metrics
- Accuracy: (TP+TN)/(P+N)
- Matthews Correlation Coefficient: 1=perfect, 0=random, -1=completely inaccurate
Algorithms Employed
- Logistic Regression (LR): Uses the sigmoid logistic equation.
- Linear Discriminant Analysis (LDA): Assumes Gaussian data with the same variance.
- k Nearest Neighbors (KNN): Predictions based on the k closest instances.
- Decision Tree (DT): Binary tree structure for predictions.
- Neural Network (NN): Models the human brain’'s decision-making.
- Naive Bayes (NB): Assumes independence between features.
- Gradient Boost (GB): Combines weak learners to create a strong learner.
Engineering Goal
Produce a machine learning model for Parkinson’s diagnosis with at least 90% accuracy and/or a Matthews Correlation Coefficient of at least 0.9. Compare algorithms and parameters to determine the best model.
Dataset Description
- Source: University of Oxford
- 195 instances: 147 Parkinson’‘s subjects, 48 without Parkinson’'s
- 22 features: Characteristics like frequency, pitch, amplitude/period of the sound wave
- 1 label: 1 for Parkinson’s, 0 for no Parkinson’s
Summary of Procedure
- Split dataset into training and validation sets.
- Train each algorithm: LR, LDA, KNN, DT, NN, NB, GB.
- Evaluate results using the validation set.
- Repeat for different training/validation splits and a rescaled dataset.
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 帕金森病 的高质量、多模态真实临床数据定制解决方案。




