IRLCov19
包含1300万条COVID-19相关推文的多语言Twitter数据集,涵盖12种印度地区语言,用于疫情分析和政策研究。
基本信息
资源简介
IRLCov19是一个大型多语言Twitter数据集,包含1300万条COVID-19相关推文,覆盖2020年2月至7月期间,专门针对印度地区语言(12种)。该数据集由印度理工学院罗凯瑞分校创建,用于研究公众对疫情的反应、政策影响分析以及疫情早期检测和监控,数据模态为文本。
下载信息
注册下载
Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。
暂未开放公开下载
Tips: 该数据集属于公开下载,应该可以免费公开下载。
免登录有偿下载
Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。
提供高速下载与技术交付服务(收技术服务费,非数据销售)
暂未开放千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。
使用方式
数据集获取
git clone https://github.com/deepakuniyaliit/Covid19IRLTDataset.git
curl -L -o repo.zip https://github.com/deepakuniyaliit/Covid19IRLTDataset/archive/refs/heads/main.zip
unzip repo.zip
源站 README 摘录(使用方式)
COVID-19 Indian Regional Languages Tweets Dataset
The repository contains a collection of Indian regional languages tweets IDs related to novel coronavirus COVID-19. The dataset contains Tweets ids starting from February 2020 to July 2020. The Twitter search API was used to gather real-time tweets that contained specific keywords in the 12 languages. The languagues for which dataset is provided are Hindi (hi), Tamil (ta), Urdu (ur), Marathi (mr), Telugu (te), Gujarati (gu), Kannada (kn), Bengali (bn), Oriya (or), Malayalam (ml), Punjabi (pa), Sindhi (sd) in the decreasing order of number of tweets. To comply with Twitter’s Terms of Service, only ids of the tweets are provided in this dataset which can be used for non-commercial research purpose only.
Data Organization
- As of Mar 26, 2021 we have tweets starting from February 01, 2020 to July 31, 2020.
- Tweet-ID files are stored month wise inside directories of Indian Regional languages which are named as ISO 2 Alphabets Code of languages.
- The Tweet-ID files contain the tweets ids where all files have similar structure iso2_Month_Date_Year.txt.
Dataset Collection
- All the tweets irrespective of any particular language were collected from February 2020 to July 2020 based on some keywords and hashtags which are provided in the file keywords.txt.
- For retrieving, the full object of the tweet consider the following tools Hydrator and twarc.
Licensing
This dataset is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License (CC BY-NC-SA 4.0).By using this dataset , you agree to the terms of the LICENSE, and to all Twitter’s Terms of Service, and cite our papers:
- IRLCov19: A Large COVID-19 Multilingual Twitter Dataset of Indian Regional Languages
- Dense Vector Embedding Based Approach to Identify Prominent Disseminators From Twitter Data Amid COVID-19 Outbreak
Contact
If you have any feedback, suggestions or queries, please do reach out to deepakuniyal(AT)geu(dot)ac(dot)in
数据加载示例(表格/文本类)
import pandas as pd, glob, os
files = (glob.glob(os.path.join(path, "**", "*.csv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.tsv"), recursive=True)
+ glob.glob(os.path.join(path, "**", "*.xlsx"), recursive=True))
print("数据文件:", files)
df = pd.read_csv(files[0])
print(df.shape); print(df.columns.tolist()); print(df.head(3))
精度瓶颈?数据缺失?
当前公开数据无法满足您的算法精度?千方提供针对 新型冠状病毒肺炎 的高质量、多模态真实临床数据定制解决方案。




