抗SARS-CoV-2基准数据集
基于ChEMBL抗SARS-CoV-2生物活性数据构建的基准数据集,用于机器学习预测抗SARS-CoV-2活性。
基本信息
资源简介
该数据集从ChEMBL数据库收集抗SARS-CoV-2生物活性数据,构建基准数据集,用于通过机器学习方法预测抗SARS-CoV-2活性,以发现新的抗SARS-CoV-2化合物和药用植物。数据模态为表格,包含化合物结构和活性标签。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/daishaoxing/anti-SARS-CoV-2.git
curl -L -o repo.zip https://github.com/daishaoxing/anti-SARS-CoV-2/archive/refs/heads/main.zip
unzip repo.zip
源站 README 摘录(使用方式)
anti-SARS-CoV-2
The code and associated dataset for the study of “In silico identification of anti-SARS-CoV-2 medicinal plants using cheminformatics and machine learning”
Abstract:
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the causative pathogen of COVID-19, is spreading rapidly and has caused hundreds of millions of infections and millions of deaths worldwide. Due to the lack of specific vaccines and effective treatments for COVID-19, there is an urgent need to identify effective drugs. Traditional Chinese medicine (TCM) is a valuable resource for identifying novel anti-SARS-CoV-2 drugs based on the important contribution of TCM and its potential benefits in COVID-19 treatment. Herein, we aimed to discover novel an-ti-SARS-CoV-2 compounds and medicinal plants from TCM by establishing a prediction method of anti-SARS-CoV-2 activity using machine learning methods. We firstly constructed a benchmark dataset from anti-SARS-CoV-2 bioactivity data collected from the ChEMBL database. Then, we established random forest (RF) and support vector machine (SVM) models that both achieved satisfactorily predictive performance with AUC values of 0.90. By using this method, a total of 1013 active anti-SARS-CoV-2 compounds were predicted from the TCMSP database. Among these compounds, six compounds with highly potent activity were confirmed in the anti-SARS-CoV-2 experiments. The molecular fingerprint similarity analysis revealed that only 24 of the 1013 compounds have high similarity to the FDA-approved antiviral drugs, indicating that most of the compounds were structurally novel. Based on the predicted anti-SARS-CoV-2 compounds, we identified 74 anti-SARS-CoV-2 medicinal plants through enrichment analysis. These identified plants are widely distributed in 68 genera and 43 families. In summary, this study provided several medicinal plants with potential anti-SARS-CoV-2 activity, which offer an attractive starting point and a broader scope to mine for potentially novel anti-SARS-CoV-2 drugs.
Introduction
This repository contains an anti-SARS-CoV-2 compound predictor constructed based on machine learning methods for screening novel anti-SARS-CoV-2 drugs from large compound database.
Requirement
the predictor is a program developing with Python. Pybel, a python wrapper of Openbabel, was used to deal with compounds and generate molecular fingerprints fo
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。




