结核分枝杆菌耐药性数据集
基于SNP和基因数据预测结核分枝杆菌药物抗性的数据集,包含样本药物敏感性和基因突变信息。
基本信息
模态
表格
创建/更新时间
2021-09-28
资源简介
该数据集包含结核分枝杆菌的SNP和基因数据,用于预测药物抗性。包括多个文件,如AllLabels.csv、SNPList.csv等,详细记录了样本对不同药物的敏感性/抗性状态及基因突变信息。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/AmirHoseinSafari/M.tuberculosis-dataset-for-drug-resistant.git
curl -L -o repo.zip https://github.com/AmirHoseinSafari/M.tuberculosis-dataset-for-drug-resistant/archive/refs/heads/main.zip
unzip repo.zip
源站 README 摘录(使用方式)
M.tuberculosis dataset for drug resistant
The SNP and gene datasets of M. Tuberculosis for drug resistance prediction.
Here is a brief description of each file:
AllLabels.csvcontains the susceptibility/resistance status (susceptibility:0 and resistance:1) for each sample isolate to 12 different drugs.SNPList.csvcontains the list of all loci on the MTB genome where a mutation was detected using the variant calling tools, based on the reference genome provided here.SNP_data_part*.zipcontains csv files with the binary SNPs. The csv files are concatenated using loading_data package (refer to this repo).gene_data.csv.zipcontians a csv file that summarizes the SNPs based on the gene that they fall into to form a matrix that contains a single feature for each gene of each sample isolate.iso_list.csva list of all isolates IDs used in the training data.sparsetableFeb27.npzThe binary SNP file in npz format for ease of use.
For understanding how to load and use this data please visit the LRCN-drug-resistance repository, especially the loading_data section.
Citation
If you found the content of this repository useful, please cite us:
https://dl.acm.org/doi/abs/10.1145/3459930.3469534
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
完整仓库:github.com/AmirHoseinSafari/M.tuberculosis-dataset-for-drug-resistant
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 结核病 的高质量、多模态真实临床数据定制解决方案。




