多标签传染病新闻事件语料库
包含多标签传染病新闻事件文本片段及标注,用于传染病事件检测与信息抽取。
基本信息
资源简介
该数据集包含多标签传染病新闻事件语料库,涵盖细粒度和粗粒度的文本片段及其标签,并附有标注指南。数据模态为文本,任务为多标签事件分类,用于传染病新闻事件的信息抽取和检索研究。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/jpiskorski/infectious-diseases-events.git
curl -L -o repo.zip https://github.com/jpiskorski/infectious-diseases-events/archive/refs/heads/main.zip
unzip repo.zip
源站 README 摘录(使用方式)
[AVAILABLE from 2 APRIL 2023]
This repository contains the Multi-label Infectious Disease News Event Corpus
mentioned in the following paper:
Multi-label Infectious Disease News Event Corpus.
Jakub Piskorski, Nicolas Stefanovitch, Brian Doherty, Jens P. Linge,
Sopho Kharazi, Jas Mantero, Guillaume Jacquet, Alessio Spadaro and Giulia Teodori.
In Proceedings of Text2Story 2023: Sixth International Workshop on Narrative Extraction
from Texts held in conjunction with the 45th European Conference on Information Retrieval,
Dublin, Ireland, 2023.
Please cite using this reference: https://github.com/jpiskorski/infectious-diseases-events/blob/main/reference.bib
The archive contains three files:
- infectious_diseases_finegrained_grained.txt
The text snippets labelled with fine-grained event types.
The text snippets and the labels are separated by tabs. - infectious_diseases_coarse_grained.txt
The text snippets labelled with coarse-grained event types.
The text snippets and the labels are separated by tabs. - Annotation_guidelines.pdf
The draft version of the annotation guidelines used by the annotators. A new complete version will be released soon.
NOTE: new updated version 1.1 available since 25 May 2023 (small fixes related to inconsistent labels and removing some redundant entries)
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 感染性疾病 的高质量、多模态真实临床数据定制解决方案。




