FairGlucose

FairGlucose是一个包含300名患者、按年龄/性别/糖尿病类型分层平衡的连续血糖监测公平性基准数据集,用于评估2小时血糖预测中的亚组公平性。

约翰霍普金斯大学; Welldoc公司约翰霍普金斯大学; Welldoc公司
arXiv
2026-08-19 更新
浏览 13
时间序列糖尿病时间序列

基本信息

模态
时间序列
创建/更新时间
2026-08-19

资源简介

FairGlucose是一个用于连续血糖监测(CGM)公平性评估的基准数据集,由约翰霍普金斯大学和Welldoc公司共同构建。数据集包含300名患者,按年龄(18-39/40-64/65+)、性别(男/女)和糖尿病类型(1型/2型)的12个交叉分层均匀平衡,每个亚组25人。数据集包含132,480个标准化预测样本(24小时输入窗口+8小时输出窗口)和81名患者记录的3,945个带时间戳的行为事件(餐饮、运动、用药)。该数据集旨在解决现有CGM基准数据人口统计不平衡的问题,用于评估2小时血糖预测任务中的亚组公平性,揭示群体层面验证指标可能掩盖的显著亚组差异。数据模态为时间序列,任务为2小时血糖预测,用途为公平性评估。

原始链接

http://arxiv.org/abs/2608.18296v1

访问原始数据

官方服务

如需原始数据获取支持或标注服务,请联系我们。

帮我联系

下载信息

注册下载

Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。

暂未开放

公开下载

Tips: 该数据集属于公开下载,应该可以免费公开下载。

免登录

有偿下载

Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。

提供高速下载与技术交付服务(收技术服务费,非数据销售)

暂未开放

千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。

使用方式

数据集说明

FairGlucose 对应论文数据集(arXiv 预印本)。

数据获取指引

  1. 打开论文页面获取作者与项目信息:https://arxiv.org/abs/2608.18296v1
  2. 论文 Data Availability / Code Availability 章节标注了数据实际托管位置;
  3. 获取到实际数据链接后,按对应平台标准方式下载。

论文摘要:Abstract:As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.

论文页面:https://arxiv.org/abs/2608.18296v1

精度瓶颈?数据缺失?

当前公开数据无法满足您的算法精度?千方提供针对 糖尿病 的高质量、多模态真实临床数据定制解决方案。

获取专属数据定制方案