FakenewsBR: A Unified Portuguese Misinformation Corpus for Reproducible Data Integration Research

Resumo


The growing impact of misinformation in digital environments has intensified the demand for structured and reusable datasets to support empirical research. This paper presents FakenewsBR, a consolidated Portuguese-language fake news resource comprising 51,205 instances integrated from six previously curated corpora rather than newly collected raw data. Through a reproducible pipeline, heterogeneous sources were harmonized under a unified schema, subjected to controlled cleaning procedures, and processed using near-duplicate analysis to flag highly similar records for manual review. The final dataset preserves provenance information, spans a temporal window from 2005 to 2022, and includes fact-check metadata for 18.18% of instances. Rather than proposing new detection models, the contribution focuses on systematic integration, structural normalization, and transparent documentation to facilitate reuse in machine learning, database research, and misinformation analysis. The dataset and its construction pipeline are publicly available at: https://github.com/AKCIT-FN/fakenews-data.

Palavras-chave: Fake news, Misinformation, Portuguese-language corpus, Dataset integration

Referências

Almada, F. L. N., Mariano, K. D. P., Dutra, M. A., da Silva Monteiro, V. E., Gomes, J. R. S., Filho, A. R. G., and da Silva Soares, A. (2025). Akcit-fn at checkthat! 2025: Switching fine-tuned slms and llm prompting for multilingual claim normalization.

Bender, E. M. and Friedman, B. (2018). Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.

Broder, A. Z. (2000). Identifying and filtering near-duplicate documents. In Giancarlo, R. and Sankoff, D., editors, Combinatorial Pattern Matching, pages 1–10. Springer Berlin Heidelberg, Berlin, Heidelberg.

Buneman, P., Khanna, S., and Wang-Chiew, T. (2001). Why and where: A characterization of data provenance. In Van den Bussche, J. and Vianu, V., editors, Database Theory — ICDT 2001, pages 316–330, Berlin, Heidelberg. Springer Berlin Heidelberg.

Cabral, L., Monteiro, J. M., Franco da Silva, J. W., Mattos, C. L., and Mourão, P. J. C. (2021). Fakewhastapp.br: Nlp and machine learning techniques for misinformation detection in brazilian portuguese whatsapp messages. In Proceedings of the 23rd International Conference on Enterprise Information Systems - Volume 1: ICEIS, pages 63–74. INSTICC, SciTePress.

Couto, J., Pimenta, B., de Araújo, I. M., Assis, S., Reis, J. C. S., da Silva, A. P., Almeida, J., and Benevenuto, F. (2021). Central de fatos: Um repositório de checagens de fatos. In Anais do III Dataset Showcase Workshop, pages 128–137, Porto Alegre, RS, Brasil. SBC.

Gomes, J., Neto, V., Barbosa, J., and de Lima, E. (2023). A rapid tertiary review at the fake news domain. In Anais da XI Escola Regional de Informática de Goiás, Porto Alegre, RS, Brasil. SBC.

Gomes, J. R. S. (2025). Verificação semi-automática de fatos em português: enriquecimento de corpus via busca e extração de alegação. Dissertação (mestrado em ciência da computação), Universidade Federal de Goiás, Goiânia, Brasil. Instituto de Informática.

Gôlo, M., Mori, A., Oliveira, W., Barbosa, J., Graciano-Neto, V., Lima, E., and Marcacini, R. (2024). On the use of large language models to detect brazilian politics fake news. In Anais do XXI Encontro Nacional de Inteligência Artificial e Computacional, pages 1–12, Porto Alegre, RS, Brasil. SBC.

Hovy, D. and Spruit, S. L. (2016). The social impact of natural language processing. In Erk, K. and Smith, N. A., editors, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 591–598, Berlin, Germany. Association for Computational Linguistics.

Lazer, D. M. J., Baum, M. A., Benkler, Y., Berinsky, A. J., Greenhill, K. M., Menczer, F., Metzger, M. J., Nyhan, B., Pennycook, G., Rothschild, D., Schudson, M., Sloman, S. A., Sunstein, C. R., Thorson, E. A., Watts, D. J., and Zittrain, J. L. (2018). The science of fake news. Science, 359(6380):1094–1096.

Martins, A. D., Cabral, L., Mourão, P., de Sá, I., Monteiro, J., and Machado, J. (2021). Covid19.br: A dataset of misinformation about covid-19 in brazilian portuguese whatsapp messages. In Anais do III Dataset Showcase Workshop, pages 138–147, Porto Alegre, RS, Brasil. SBC.

Monteiro, R. A., Santos, R. L. S., Pardo, T. A. S., de Almeida, T. A., Ruiz, E. E. S., and Vale, O. A. (2018). Contributions to the study of fake news in portuguese: New corpus and automatic detection results. In Computational Processing of the Portuguese Language: 13th International Conference, PROPOR 2018, Canela, Brazil, September 24–26, 2018, Proceedings, page 324–334, Berlin, Heidelberg. Springer-Verlag.

Nielsen, D. S. and McConville, R. (2022). Mumin: A large-scale multilingual multimodal fact-checked misinformation social network dataset. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’22, page 3141–3153, New York, NY, USA. Association for Computing Machinery.

Ribeiro, M., Calais, P., Santos, Y., Almeida, V., and Meira Jr., W. (2018). Characterizing and detecting hateful users on twitter. Proceedings of the International AAAI Conference on Web and Social Media, 12(1).

Santos, Y., Silva, M., and Reis, J. C. S. (2023). Imagefactck.br: Repositório de imagens para a detecção de desinformação disseminada em plataformas digitais. In Anais do V Dataset Showcase Workshop, pages 87–98, Porto Alegre, RS, Brasil. SBC.

Selau, F. (2021). Fakes news in portuguese. [link]. Accessed: April 14, 2026.

Shu, K., Mahudeswaran, D., Wang, S., Lee, D., and Liu, H. (2020). Fakenewsnet: A data repository with news content, social context, and spatiotemporal information for studying fake news on social media. Big Data, 8(3):171–188. PMID: 32491943.

Shu, K., Sliva, A., Wang, S., Tang, J., and Liu, H. (2017). Fake news detection on social media: A data mining perspective. SIGKDD Explor. Newsl., 19(1):22–36.

Silva, R. M., Santos, R. L., Almeida, T. A., and Pardo, T. A. (2020). Towards automatically filtering fake news in portuguese. Expert Syst. Appl., 146(C).

Wang, W. Y. (2017). “liar, liar pants on fire”: A new benchmark dataset for fake news detection. In Barzilay, R. and Kan, M.-Y., editors, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 422–426, Vancouver, Canada. Association for Computational Linguistics.
Publicado
08/09/2026
MARIANO, Kauan Divino Pouso; ALMADA, Fabrycio Leite Nakano; DUTRA, Maykon Adriell; MONTEIRO, Victor Emanuel da Silva; GOMES, Juliana Resplande Sant’Anna; GALVÃO FILHO, Arlindo Rodrigues; SOARES, Anderson da Silva. FakenewsBR: A Unified Portuguese Misinformation Corpus for Reproducible Data Integration Research. In: DATASET SHOWCASE WORKSHOP (DSW), 8. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 24-35. DOI: https://doi.org/10.5753/dsw.2026.249457.