HISTAI-结直肠-b2
HISTAI-结直肠-b2是一个大型全切片图像数据集,用于计算病理学中的结直肠癌AI诊断研究。
基本信息
资源简介
该数据集是一个大型全切片图像数据集,专注于结直肠病理学,包含大量标注的全切片图像,可用于计算病理学中的分类、检测等任务。数据由HISTAI团队收集并开放,用于推动结直肠癌的AI诊断研究。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取(ModelScope)
方式一:MsDataset(Python)
# 前置依赖: pip install modelscope
from modelscope.msdatasets import MsDataset
ds = MsDataset.load("histai/HISTAI-colorectal-b2", subset_name="default", split="train")
print(ds)
方式二:CLI(命令行)
pip install modelscope
modelscope download dataset histai/HISTAI-colorectal-b2
数据集卡片摘录(源站)
Dear researchers and engineers, you’‘re accessing a dataset that would cost millions of dollars to build and took millions of nerves to negotiate favorable terms for its use. Your support, by liking the repositories and upvoting the collection, costs nothing but gives us valuable motivation to continue our contributions to the community. We reserve the right not to approve the request if you don’'t support our efforts. Thank you very much for collaboration!
HISTAI Dataset
Find more information and metadata at histai/HISTAI-metadata.
📄 Paper and Citation
A detailed research paper describing the HISTAI dataset is available:
Dmitry Nechaev, Alexey Pchelnikov, Ekaterina Ivanova
📖 Read on arXiv
📚 Citation
If you use HISTAI in your research, please cite:
@misc{nechaev2025histaiopensourcelargescaleslide,
title = {HISTAI: An Open-Source, Large-Scale Whole Slide Image Dataset for Computational Pathology},
author = {Dmitry Nechaev and Alexey Pchelnikov and Ekaterina Ivanova},
year = {2025},
eprint = {2505.12120},
archivePrefix = {arXiv},
primaryClass = {eess.IV},
url = {https://arxiv.org/abs/2505.12120}
}
## 许可
cc-by-nc-4.0
## 数据加载示例(图像类)
```python
from PIL import Image
import glob, os
files = (glob.glob(os.path.join(path, "**", "*.png"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.jpg"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tif"), recursive=True))
print("图像文件数:", len(files))
img = Image.open(files[0]); print("尺寸/模式:", img.size, img.mode)
# torchvision Dataset 方式:
# from torchvision import datasets
# ds = datasets.ImageFolder(path) # 要求 子目录=类别
目录组织与标注格式以源站说明和下载后实际文件为准。
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 直肠癌 的高质量、多模态真实临床数据定制解决方案。




