Language Models and Detection Strategies for Political Disinformation in Brazil: A Comparative Applied Study

Resumo


The rapid spread of online political disinformation has increased the need for scalable detection beyond manual fact-checking. This study evaluates textual disinformation detection in Brazilian Portuguese by comparing BERTimbau-based supervised models with open-source LLMs under zero-shot, Chain-of-Thought, and RAG settings. Four public corpora were used for training, validation, and retrieval, while an external political dataset supported out-of-distribution testing. Results show that gemma2:9b in zero-shot achieved the best macro-F1 (0.950), outperforming BERTimbau+SVM (0.837). BM25-based RAG reached macro-F1 = 0.923 with grounded decisions. Overall, performance depends on inference strategy, architecture, and corpus characteristics.

Palavras-chave: Political Disinformation Detection, Large Language Models, Retrieval-Augmented Generation, Out-of-Distribution Evaluation, Brazilian Portuguese

Referências

Bernardi, A. J. B. (2021). Fake news e as eleições de 2018 no Brasil: como diminuir a desinformação? Editora Appris.

Boumber, D., Tuck, B. E., Verma, R. M., and Qachfar, F. Z. (2024). Llms for explainable few-shot deception detection. In Proceedings of the 10th ACM International Workshop on Security and Privacy Analytics, pages 37–47.

BRASIL (2019). Resolução tse n.º 23.610, de 18 de dezembro de 2019. Tribunal Superior Eleitoral, Brasília. Dispõe sobre propaganda eleitoral e condutas ilícitas em campanha.

Cabral, L., Monteiro, J., da Silva, J., Mattos, C., and Mourão, P. (2021). Fakewhatsapp.br: Nlp and machine learning techniques for misinformation detection in brazilian portuguese whatsapp messages. In ICEIS, pages 63–74.

Cantarino, F. H. S. (2024). Criação de um corpus português para auxiliar a identificação de notícias verdadeiras e falsas.

Charles, A. C., Ruback, L., and Oliveira, J. (2022). Fakepedia corpus: A flexible fake news corpus in portuguese. In International Conference on Computational Processing of the Portuguese Language, pages 37–45. Springer.

Chavarro, J. P., Carvalho, J. T., Portela, T. T., and Silva, J. C. (2023). Faketruebr: Um corpus brasileiro de notícias falsas. In Escola Regional de Banco de Dados (ERBD), pages 108–117. SBC.

Cordeiro, P. and Pinheiro, V. (2019). Um corpus de notícias falsas do twitter e verificação automática de rumores em língua portuguesa. In Proceedings of the Symposium in Information and Human Language Technology, pages 219–228.

Couto, J. M., Pimenta, B., de Araújo, I. M., Assis, S., Reis, J. C., da Silva, A. P. C., Almeida, J. M., and Benevenuto, F. (2021). Central de fatos: Um repositório de checagens de fatos. In Dataset Showcase Workshop (DSW), pages 128–137. SBC.

de Morais, J., Abonizio, H., Tavares, G., da Fonseca, A., and Barbon Jr, S. (2020). A multi-label classification system to distinguish among fake, satirical, objective and legitimate news in brazilian portuguese. iSys – Brazilian Journal of Information Systems, 13(4):126–149.

Delucis, M. M., Fraga, L., Parraga, O., Mattjie, C., Ravazio, R., Barros, R. C., and Kupssinskü, L. S. (2025). Automated fact-checking in brazilian portuguese: Resources and baselines. In Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL), pages 137–148. SBC.

Faustini, P. and Covões, T. (2019). Fake news detection using one-class classification. In 2019 8th Brazilian Conference on Intelligent Systems (BRACIS), pages 592–597. IEEE.

Garcia, G. L., Afonso, L. C., and Papa, J. P. (2022). Fakerecogna: A new brazilian corpus for fake news detection. In International Conference on Computational Processing of the Portuguese Language, pages 57–67. Springer.

Garcia, K., Shiguihara, P., and Berton, L. (2024). Breaking news: Unveiling a new dataset for portuguese news classification and comparative analysis of approaches. Plos one, 19(1):e0296929.

Gôlo, M. P. S., Mori, A. L. V., Oliveira, W. G., Barbosa, J. R., Graciano Neto, V. V., Lima, E. A. d., and Marcacini, R. M. (2024). On the use of large language models to detect brazilian politics fake news. Anais do Encontro Nacional de Inteligência Artificial e Computacional.

Gôlo, M., Caravanti, M., Rossi, R., Rezende, S., Nogueira, B., and Marcacini, R. (2021). Learning textual representations from multiple modalities to detect fake news through one-class learning. In Brazilian Symposium on Multimedia and the Web, pages 197–204.

Herculano, A., da Silva, L., Fernandes, D., and Rego, A. (2025). Uma avaliação comparativa entre o deprebertbr e modelos de linguagem para classificação de textos depressivos. In Anais do XL Simpósio Brasileiro de Bancos de Dados, pages 535–548, Porto Alegre, RS, Brasil. SBC.

Hu, B., Sheng, Q., Cao, J., Shi, Y., Li, Y., Wang, D., and Qi, P. (2024). Bad actor, good advisor: Exploring the role of large language models in fake news detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 22105–22113. Number 20.

Kansaon, D., de Freitas Melo, P., Zannettou, S., and Benevenuto, F. (2025). From fake news to real protests: Whatsapp’s role in brazilian political coordination. In Proceedings of the International AAAI Conference on Web and Social Media, volume 19, pages 1007–1020.

Li, Y., Jiang, B., Shu, K., and Liu, H. (2020). Mm-covid: A multilingual and multimodal data repository for combating covid-19 disinformation.

Liu, B., Wang, A., and Xia, C. (2025). Interpretable chinese fake news detection with chain-of-thought and in-context learning. IEEE Access.

Lopez-Joya, S., Diaz-Garcia, J. A., Ruiz, M. D., and Martin-Bautista, M. J. (2025). The blueprint of a new fact-checking system: A methodology to enrich rag systems with new generated datasets. Computers and Electrical Engineering, 128:110746.

Martins, A. D. F., Cabral, L., Mourão, P. J. C., de Sá, I. C., Monteiro, J. M., and Machado, J. (2021). Covid19.br: A dataset of misinformation about covid-19 in brazilian portuguese whatsapp messages. In Dataset Showcase Workshop (DSW), pages 138–147. SBC.

Monteiro, R. A., Santos, R., Pardo, T., de Almeida, T., Ruiz, E., and Vale, O. (2018). Contributions to the study of fake news in portuguese: New corpus and automatic detection results. In Computational Processing of the Portuguese Language, pages 324–334. Springer.

Moreno, J. A. and Bressan, G. (2019). Factck.br: A new dataset to study fake news. In Proceedings of the 25th Brazilian Symposium on Multimedia and the Web, pages 525–527. ACM.

Nielsen, D. S. and McConville, R. (2022). MuMiN: A large-scale multilingual multimodal fact-checked misinformation social network dataset. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 3141–3153. ACM.

Paiva, L., Assis, G., Amorim, A., Dias, L. G., Paes, A., and de Oliveira, D. (2025). Domínio delimitado, Ódio exposto: O uso de prompts para identificação de discurso de Ódio online com llms. In Anais do XL Simpósio Brasileiro de Bancos de Dados, pages 493–506, Porto Alegre, RS, Brasil. SBC.

Patel, A., Tiwari, A., and Ahmad, S. (2022). Fake news detection using support vector machine. In Proceedings of the 3rd International Conference on Advanced Computing and Software Engineering (ICACSE 2021), pages 34–38. SCITEPRESS.

Pires, V. B., Guerreiro, D., et al. (2024). Portuguese fake news classification with bert models. In Encontro Nacional de Inteligência Artificial e Computacional (ENIAC), pages 834–845. SBC.

Schreiber, A. (2022). Civil rights framework of the internet (bcrfi; marco civil da internet): Advance or setback? civil liability for damage derived from content generated by third party. In Personality and Data Protection Rights on the Internet: Brazilian and German Approaches, pages 241–266. Springer.

Shahi, G. K. and Nandini, D. (2020). Fakecovid – a multilingual cross-domain fact check news dataset for covid-19. In Proceedings of ICWSM.

Silva, F. R. M. d. (2020). Fakenewssetgen: um processo para construção de datasets que viabilizem a comparação entre métodos de detecção de fake news. Master’s thesis, Instituto Militar de Engenharia (IME).

Silva, R. M., Amamou, H., Ferraz, L. B. S., da Silva, F. K. A., and Avila, A. R. (2025). Fake news detection in portuguese under large language model-generated content. Journal of the Brazilian Computer Society, 31(1):1149–1166.

Vosoughi, S., Roy, D., and Aral, S. (2018). The spread of true and false news online. science, 359(6380):1146–1151.
Publicado
08/09/2026
MORI, Adriel L. V.; SANTOS, Willgnner Ferreira; LIMA, Eliomar Araújo; GRACIANO-NETO, Valdemar V.; BARBOSA, Jacson R.; CARVALHO, Rogerio R.; DA SILVA, Nádia F. F.; GALVÃO FILHO, Arlindo R.. Language Models and Detection Strategies for Political Disinformation in Brazil: A Comparative Applied Study. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 714-727. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249311.