抗癌症肽数据集

整合多个数据库的抗癌症肽,包含序列、来源、实验证据及负对照,用于癌症治疗研究。

Sadik90Sadik90
GitHub
2024-04-19 更新
浏览 23
蛋白质抗癌症肽癌症治疗

基本信息

模态
蛋白质
创建/更新时间
2024-04-19

资源简介

该数据集整合了来自多个数据库的抗癌症肽,用于癌症治疗研究。包含多样化的抗癌症肽序列、来源、实验证据和相关参考文献,以及精心挑选的非分泌蛋白负对照集。

原始链接

https://github.com/Sadik90/ACP-Dataset

访问原始数据

官方服务

如需原始数据获取支持或标注服务,请联系我们。

帮我联系

下载信息

注册下载

Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。

暂未开放

公开下载

Tips: 该数据集属于公开下载,应该可以免费公开下载。

免登录

有偿下载

Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。

提供高速下载与技术交付服务(收技术服务费,非数据销售)

暂未开放

千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。

使用方式

数据集获取

git clone https://github.com/Sadik90/ACP-Dataset.git

curl -L -o repo.zip https://github.com/Sadik90/ACP-Dataset/archive/refs/heads/main.zip
unzip repo.zip

源站 README 摘录(使用方式)

ACP-Dataset

Release of Anticancer Peptide Dataset Integrated from Multiple Databases for Cancer Therapeutics Research
Introduction:
The release of the anticancer peptide dataset represents a significant advancement in cancer therapeutics research, aimed at facilitating the development of novel peptide-based treatments for various cancer types. This dataset is meticulously curated from multiple databases, including CancerPPD, APD3, LAMP, and UniProt, with a special focus on integrating peptides from non-secretory proteins as a negative dataset. By compiling diverse sources of peptide information, this dataset offers a comprehensive resource for researchers to explore the potential of peptides in cancer therapy.
Dataset Description:
The dataset comprises a diverse collection of anticancer peptides derived from various biological sources, encompassing different cancer types and stages. Each peptide is annotated with detailed information, including its sequence, origin, experimental evidence, and relevant references. Furthermore, the dataset includes a carefully curated set of non-secretory proteins, serving as a negative dataset for comparative analysis.
Integration with Cancer Genomic Datasets:
To enhance the utility of the anticancer peptide dataset, researchers can integrate it with cancer genomic datasets such as TCGA (The Cancer Genome Atlas) and Cancer Genome databases. By correlating peptide sequences with cancer subtypes and genomic alterations, researchers can identify potential therapeutic targets and elucidate the molecular mechanisms underlying cancer progression.
Potential Applications:
The integrated dataset enables researchers to explore various applications in cancer therapeutics, including:
Targeted Therapy: By combining peptide sequences with DNA methylation data, researchers can develop targeted therapies that exploit specific methylation patterns associated with cancer subtypes.
Splicing Therapy: Integration with alternate splicing data allows for the identification of splice variants associated with cancer progression, paving the way for splicing-based therapeutic interventions.
De Novo Peptide Design: Leveraging advanced generative models such as LSTM (Long Short-Term Memory), RNN (Recurrent Neural Network), VAE (Variational Autoencoder), and GAN (Generative Adversarial Network), researchers can generate novel anticancer peptides tailore

数据加载示例(表格/文本类)

import pandas as pd, glob, os

files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
       + glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
       + glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))

完整仓库:github.com/Sadik90/ACP-Dataset

精度瓶颈?数据缺失?

当前公开数据无法满足您的算法精度?千方提供针对 癌症(总论) 的高质量、多模态真实临床数据定制解决方案。

获取专属数据定制方案