• Home
  • About
  • Repositories
  • Search
  • Web API
  • Feedback
<< Go Back

Metadata

Name
On the Role of Images for Analyzing Claims in Social Media
Repository
ZENODO
Identifier
doi:10.5281/zenodo.4592249
Description
This is a multimodal dataset used in the paper &quot;On the Role of Images for Analyzing Claims in Social Media&quot;, accepted at CLEOPATRA-2021 (2nd International Workshop on Cross-lingual Event-centric Open Analytics), co-located with The Web Conference 2021.

The four datasets are curated for two different tasks that broadly come under fake news detection. Originally, the datasets were released as part of challenges or papers for text-based NLP tasks and are further extended here with corresponding images.

1. clef_en and clef_ar are English and Arabic Twitter datasets for claim check-worthiness detection released in CLEF CheckThat! 2020 Barr&oacute;n-Cedeno et al. [1].
2. lesa is an English Twitter dataset for claim detection released by Gupta et al.[2]
3. mediaeval is an English Twitter dataset for conspiracy detection released in MediaEval 2020 Workshop by Pogorelov et al.[3]

The dataset details like data curation and annotation process can be found in the cited papers.

Datasets released here with corresponding images are relatively smaller than the original text-based tweets. The data statistics are as follows:
1. clef_en: 281
2. clef_ar: 2571
3. lesa: 1395
4. mediaeval: 1724

Each folder has two sub-folders and a json file data.json that consists of crawled tweets. Two sub-folders are:
1. images: This Contains crawled images with the same name as tweet-id in data.json.
2. splits: This contains 5-fold splits used for training and evaluation in our paper. Each file in this folder is a csv with two columns &lt;tweet-id, label&gt;.

Code for the paper: https://github.com/cleopatra-itn/image_text_claim_detection

If you find the dataset and the paper useful, please cite our paper and the corresponding dataset papers[1,2,3]
Cheema, Gullal S., et al. &quot;On the Role of Images for Analyzing Claims in Social Media&quot; 2nd International Workshop on Cross-lingual Event-centric Open Analytics (CLEOPATRA) co-located with The Web Conf 2021.

&nbsp;

[1] Barr&oacute;n-Cedeno, Alberto, et al. &quot;Overview of CheckThat! 2020: Automatic identification and verification of claims in social media.&quot; International Conference of the Cross-Language Evaluation Forum for European Languages. Springer, Cham, 2020.
[2] Gupta, Shreya, et al. &quot;LESA: Linguistic Encapsulation and Semantic Amalgamation Based Generalised Claim Detection from Online Content.&quot; arXiv preprint arXiv:2101.11891 (2021).
[3] Pogorelov, Konstantin, et al. &quot;FakeNews: Corona Virus and 5G Conspiracy Task at MediaEval 2020.&quot; MediaEval 2020 Workshop. 2020.

&nbsp;
Data or Study Types
multiple
Source Organization
Unknown
Access Conditions
available
Year
2021
Access Hyperlink
https://doi.org/10.5281/zenodo.4592249

Distributions

  • Encoding Format: HTML ; URL: https://doi.org/10.5281/zenodo.4592249
This project was funded in part by grant U24AI117966 from the NIH National Institute of Allergy and Infectious Diseases as part of the Big Data to Knowledge program. We thank all members of the bioCADDIE community for their valuable input on the overall project.