糖尿病数据集(皮马印第安人)
包含768个Pima印第安女性样本的糖尿病预测数据集,8个数值特征,二分类标签。
基本信息
资源简介
该数据集源于美国国家糖尿病-消化-肾脏疾病研究所,包含768个21岁以上Pima印第安女性的医疗记录,拥有8个数值型自变量(如怀孕次数、血糖浓度等),目标变量为糖尿病检测结果(二分类:阳性/阴性)。数据模态为表格,主要用于二分类预测任务。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/SerdarTafrali/Machine_Learning_Pipeline_on_Diabetes_Dataset.git
curl -L -o repo.zip https://github.com/SerdarTafrali/Machine_Learning_Pipeline_on_Diabetes_Dataset/archive/refs/heads/master.zip
unzip repo.zip
源站 README 摘录(使用方式)
Machine Learning Pipeline on Diabetes Dataset
Business Problem:
Developing a machine learning model that can predict whether people have diabetes when their characteristics are specified.
Dataset Story:
The dataset is part of the large dataset held at the National Institutes of Diabetes-Digestive-Kidney Diseases in the USA. Data used for diabetes research on Pima Indian women aged 21 and over living in Phoenix, the 5th largest city of the State of Arizona in the USA. It consists of 768 observations and 8 numerical independent variables. The target variable is specified as “outcome”; 1 indicates positive diabetes test result, 0 indicates negative.
Variables:
Pregnancies: Number of pregnancies Glucose: Glucose. BloodPressure: Blood pressure. SkinThickness: Skin Thickness Insulin: Insulin. BMI: Body mass index. DiabetesPedigreeFunction: A function that calculates our probability of having diabetes based on our ancestry. Age: Age (years) Outcome: Information whether the person has diabetes or not. Have the disease (1) or not (0))
Project Stages:
- Exploratory Data Analysis
- Data Preprocessing
- Model & Prediction
- Model Evaluation
- Model Validation: Holdout
- Model Validation: 10-Fold Cross Validation
- Prediction for A New Observation
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
完整仓库:github.com/SerdarTafrali/Machine_Learning_Pipeline_on_Diabetes_Dataset
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 糖尿病 的高质量、多模态真实临床数据定制解决方案。




