Histo-DD
Histo-DD是针对组织病理学图像的数据集蒸馏算法,通过生成合成样本减少训练成本,提升分类效率。
基本信息
资源简介
Histo-DD是一个针对组织病理学图像的数据集蒸馏算法,由澳大利亚健康创新研究所开发。该算法通过整合染色标准化和模型增强技术,将大型数据集压缩成一组合成样本,以提高训练效率和简化下游应用。主要应用于组织病理学图像的分类任务,通过生成合成样本减少训练所需的大量补丁,同时保留区分性信息,显著降低训练成本。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集说明
Histo-DD 对应论文数据集(arXiv 预印本)。
数据获取指引
- 打开论文页面获取作者与项目信息:https://arxiv.org/abs/2408.09709v1
- 论文 Data Availability / Code Availability 章节标注了数据实际托管位置;
- 获取到实际数据链接后,按对应平台标准方式下载。
论文摘要:Abstract:Deep neural networks (DNNs) have exhibited remarkable success in the field of histopathology image analysis. On the other hand, the contemporary trend of employing large models and extensive datasets has underscored the significance of dataset distillation, which involves compressing large-scale datasets into a condensed set of synthetic samples, offering distinct advantages in improving training efficiency and streamlining downstream applications. In this work, we introduce a novel dataset distillation algorithm tailored for histopathology image datasets (Histo-DD), which integrates stain normalisation and model augmentation into the distillation progress. Such integration can substantially enhance the compatibility with histopathology images that are often characterised by high colour heterogeneity. We conduct a comprehensive evaluation of the effectiveness of the proposed algorithm and the generated histopathology samples in both patch-level and slide-level classification tasks. The experimental results, carried out on three publicly available WSI datasets, including Camelyon16, TCGA-IDH, and UniToPath, demonstrate that the proposed Histo-DD can generate more informative synthetic patches than previous coreset selection and patch sampling methods. Moreover, the synthetic samples can preserve discriminative information, substantially reduce training efforts, and exhibit architecture-agnostic properties. These advantages indicate that synthetic samples can serve as an alternative to large-scale datasets.
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 癌症(总论) 的高质量、多模态真实临床数据定制解决方案。




