皮马印第安人糖尿病数据集

包含768例Pima印第安女性临床特征,用于糖尿病二分类预测的表格数据集。

cnachtegcnachteg
GitHub
2019-02-05 更新
浏览 8
表格糖尿病表格

基本信息

模态
表格
创建/更新时间
2019-02-05

资源简介

该数据集包含768例Pima印第安女性的8个临床特征(怀孕次数、血糖浓度、血压、皮肤厚度、胰岛素、BMI、糖尿病遗传函数、年龄)以及糖尿病诊断结果,用于二分类预测糖尿病。数据来源于国家糖尿病和消化及肾脏疾病研究所,由Vincent Sigillito捐赠。

原始链接

https://github.com/cnachteg/diabetes_dataset

访问原始数据

官方服务

如需原始数据获取支持或标注服务,请联系我们。

帮我联系

下载信息

注册下载

Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。

暂未开放

公开下载

Tips: 该数据集属于公开下载,应该可以免费公开下载。

免登录

有偿下载

Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。

提供高速下载与技术交付服务(收技术服务费,非数据销售)

暂未开放

千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。

使用方式

数据集获取

git clone https://github.com/cnachteg/diabetes_dataset.git

curl -L -o repo.zip https://github.com/cnachteg/diabetes_dataset/archive/refs/heads/master.zip
unzip repo.zip

源站 README 摘录(使用方式)

diabetes_dataset

The resources for this dataset can be found at https://www.openml.org/d/37
Author: Vincent Sigillito
Source: Obtained from UCI
Please cite: UCI citation policy

  1. Title: Pima Indians Diabetes Database
  2. Sources:
    (a) Original owners: National Institute of Diabetes and Digestive and
    Kidney Diseases
    (b) Donor of database: Vincent Sigillito (vgs@aplcen.apl.jhu.edu)
    Research Center, RMI Group Leader
    Applied Physics Laboratory
    The Johns Hopkins University
    Johns Hopkins Road
    Laurel, MD 20707
    (301) 953-6231
    © Date received: 9 May 1990
  3. Past Usage:
  4. Smith,~J.~W., Everhart,~J.~E., Dickson,~W.~C., Knowler,~W.~C., &
    Johannes,~R.~S. (1988). Using the ADAP learning algorithm to forecast
    the onset of diabetes mellitus. In {it Proceedings of the Symposium
    on Computer Applications and Medical Care} (pp. 261265). IEEE
    Computer Society Press.
    The diagnostic, binary-valued variable investigated is whether the
    patient shows signs of diabetes according to World Health Organization
    criteria (i.e., if the 2 hour post-load plasma glucose was at least
    200 mg/dl at any survey examination or if found during routine medical
    care). The population lives near Phoenix, Arizona, USA.
    Results: Their ADAP algorithm makes a real-valued prediction between
    0 and 1. This was transformed into a binary decision using a cutoff of
    0.448. Using 576 training instances, the sensitivity and specificity
    of their algorithm was 76% on the remaining 192 instances.
  5. Relevant Information:
    Several constraints were placed on the selection of these instances from
    a larger database. In particular, all patients here are females at
    least 21 years old of Pima Indian heritage. ADAP is an adaptive learning
    routine that generates and executes digital analogs of perceptron-like
    devices. It is a unique algorithm; see the paper for details.
  6. Number of Instances: 768
  7. Number of Attributes: 8 plus class
  8. For Each Attribute: (all numeric-valued)
  9. Number of times pregnant
  10. Plasma glucose concentration a 2 hours in an oral glucose tolerance test
  11. Diastolic blood pressure (mm Hg)
  12. Triceps skin fold thickness (mm)
  13. 2-Hour serum insulin (mu U/ml)
  14. Body mass index (weight in kg/(height in m)^2)
  15. Diabetes pedigree function

数据加载示例(表格/文本类)

import pandas as pd, glob, os

files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
       + glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
       + glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))

完整仓库:github.com/cnachteg/diabetes_dataset

精度瓶颈?数据缺失?

当前公开数据无法满足您的算法精度?千方提供针对 糖尿病 的高质量、多模态真实临床数据定制解决方案。

获取专属数据定制方案