药物组合数据集
包含手动整理的已发表抗癌药物组合试验数据,来自PubMed,以CSV表格形式提供。
基本信息
资源简介
该数据集包含手动整理的已发表抗癌药物组合试验数据,数据来源于PubMed检索的出版物,以CSV文件形式提供,包括试验信息、药物信息等,主要用于抗癌药物组合研究。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/langli-lab/drugcombo-data.git
curl -L -o repo.zip https://github.com/langli-lab/drugcombo-data/archive/refs/heads/main.zip
unzip repo.zip
源站 README 摘录(使用方式)
Drug Combo dataset
This repository contains the manual curation of published Phase I anticancer drug trials for Drug Combo.
Data Source
Publications from PubMed search.
Data curation guideline
Data Curation Guideline depicts how the data was curated from the publications. In this guideline, some examples are provided to demonstrate where the required data elements can be found in the papers. For example, the dose levels can be found in the figures or some paragraphs.
Data Files
| File Name | Description |
|---|---|
| trial.csv | Some general information about the trials, including PMID, Trial Registed number, cancer subtypes, trial recruitment criteria. |
| drug.csv | The drug used in the trial, inclduing name, formula, and administration route. **Please note: There is no information about doing it in the file." |
| schema.csv | The statistical design of the trial |
| dlt_definition.csv | How dose-limiting toxicity (DLT) is defined in the trials |
| dose_level.csv | Detailed designed dosing information |
| mtd.csv | Maximum Tolerated Dose (MTD) |
| observed_dlt.csv | Reported DLTs in the publications |
Data Dictionary
trial.csv
| Column | Description |
|---|---|
| PMID | Pu |
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 癌症(总论) 的高质量、多模态真实临床数据定制解决方案。




