R-Super
R-Super是一个多肿瘤早期检测数据集,包含10万+份CT报告与图像,用于训练分割模型,降低对人工掩膜的需求。
基本信息
资源简介
R-Super是一个用于多肿瘤早期检测的人工智能训练数据集,包含101,654份与CT扫描图像关联的医学报告,描述了肿瘤的大小、数量、位置和衰减等信息,用于训练分割模型,降低对人工肿瘤掩膜的需求。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/MrGiovanni/R-Super.git
curl -L -o repo.zip https://github.com/MrGiovanni/R-Super/archive/refs/heads/main.zip
unzip repo.zip
源站 README 摘录(使用方式)
<h1 align=“center”>R-Super: Learning Segmentation from Radiology Reports</h1>
<div align=“center”>
Subscribe us: https://groups.google.com/u/2/g/bodymaps
<p align=“center”>
<img src=“documents/r_super_pdac-8.gif” width=“600”/>
</p>
Abdominal CT datasets have dozens to a couple thousand tumor masks. In contrast, hospitals and new public datasets have tens/hundreds of thousands of tumor CTs with radiology reports. Thus, we ask: how can radiology reports improve tumor segmentation?
We present R-Super, a training strategy that transforms radiology reports (text) into direct (per-voxel) supervision for tumor segmentation AI. Before training, we use LLM to extract tumor information from radiology reports. Then, R-Super introduces new loss functions (Volume Loss & Ball Loss), which use this extracted information to teach the AI to segment tumors that are coherent with reports, in terms of tumor count, diameters, and locations. R-Super can train AI on large-scale CT-Report datasets (e.g., Merlin) plus small or large CT-Mask datasets (e.g., AbdomenAtlas, PanTS). In comparison to traditional training with masks only, using R-Super to train with masks and reports improved the performance of tumor segmentation AI by up to +16% in sensitivity, F1, AUC, DSC and NSD.
<p align=“center”>
<img src=“documents/rsuper_abstract.png” width=“600”/>
</p>
[!NOTE]
Extension: Learning from 100,000 CT Scans and Reports, to Segment 7 Understudied Tumor Types
We have assembled more than 100,000 CT-Report pairs and trained R-Super to detect 7 tumor types missing from public CT-Mask datasets. Our preprint is available here.
Papers
<b>Learning Segmentation from Radiology Reports</b> <br/>
Pedro R. A. S. Bassi, Wenxuan Li, Jieneng Chen, Zheren Zhu, Tianyu Lin, Sergio Decherchi, Andrea Cavalli, [Kang Wang](https://radiology.ucsf
数据加载示例(图像类)
from PIL import Image
import glob, os
files = (glob.glob(os.path.join(path, "**", "*.png"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.jpg"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tif"), recursive=True))
print("图像文件数:", len(files))
img = Image.open(files[0]); print("尺寸/模式:", img.size, img.mode)
# torchvision Dataset 方式:
# from torchvision import datasets
# ds = datasets.ImageFolder(path) # 要求 子目录=类别
目录组织与标注格式以源站说明和下载后实际文件为准。
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 癌症(总论) 的高质量、多模态真实临床数据定制解决方案。




