Comparative Analysis of Language Models in the Understanding of Short Videos: A Multimodal Approach for Disinformation Detection

  • Adriel L. V. Mori Universidade Federal de Goiás (UFG) / Advanced Knowledge Center for Immersive Technologies (AKCIT) https://orcid.org/0000-0002-7794-8604
  • Eliomar Araújo Lima Universidade Federal de Goiás (UFG)
  • Valdemar V. Graciano-Neto Universidade Federal de Goiás (UFG) https://orcid.org/0000-0003-2190-5477
  • Jacson R. Barbosa Universidade Federal de Goiás (UFG)
  • Rogerio Rodrigues Carvalho Universidade Federal de Goiás (UFG) https://orcid.org/0009-0000-5889-4869
  • Arlindo R. Galvão Filho Universidade Federal de Goiás (UFG) / Advanced Knowledge Center for Immersive Technologies (AKCIT)

Resumo


Short-form videos are increasingly used to spread political disinformation, but multimodal benchmarks for Brazilian Portuguese remain limited. This study evaluates whether structured multimodal evidence improves zero-shot LLM-based detection in Brazilian political videos. We combine CLIP keyframe summarization, Whisper transcription, and pyannote-audio diarization to build prompts from auditory and visual evidence. We benchmark 12 LLMs on 119 videos from the 2022–2024 Brazilian electoral cycles, comparing audio-only and multimodal settings for veracity and taxonomy classification. Results show that gemma2-9b achieves the best veracity performance (macro-F1 = 0.764), while gemma3-12b leads taxonomy classification (macro-F1 = 0.518).

Palavras-chave: Multimodal Data Integration

Referências

Apostolidis, E., Adamantidou, E., Metsai, A. I., Mezaris, V., and Patras, I. (2021). Video summarization using deep neural networks: A survey. Proceedings of the IEEE, 109(11):1838–1863.

Atrey, P. K., Hossain, M. A., El Saddik, A., and Kankanhalli, M. S. (2010). Multimodal fusion for multimedia analysis: a survey. Multimedia systems, 16:345–379.

Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J. (2025). Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923.

Baltrušaitis, T., Ahuja, C., and Morency, L.-P. (2018). Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2):423–443.

Bouchey, B., Castek, J., and Thygeson, J. (2021). Multimodal learning. Innovative learning environments in STEM higher education: Opportunities, challenges, and looking forward, pages 35–54.

Bu, Y., Sheng, Q., Cao, J., Qi, P., Wang, D., and Li, J. (2024). FakingRecipe: Detecting fake news on short video platforms from the perspective of creative process. In Proceedings of the 32nd ACM International Conference on Multimedia (MM ’24), pages 1351–1360. Accessed: 2026-03-28.

Dvornik, M., Hadji, I., Derpanis, K. G., Garg, A., and Jepson, A. (2021). Drop-dtw: Aligning common signal between sequences while dropping outliers. Advances in Neural Information Processing Systems, 34:13782–13793.

Gôlo, M. P. S., Mori, A. L. V., Oliveira, W. G., Barbosa, J. R., Graciano Neto, V. V., Lima, E. A. d., and Marcacini, R. M. (2024). On the use of large language models to detect brazilian politics fake news. Anais do Encontro Nacional de Inteligência Artificial e Computacional (ENIAC).

Hua, J., Cui, X., Li, X., Tang, K., and Zhu, P. (2023). Multimodal fake news detection through data augmentation-based contrastive learning. Applied Soft Computing, 136:110125.

Jing, J., Wu, H., Sun, J., Fang, X., and Zhang, H. (2023). Multimodal fake news detection via progressive fusion networks. Information Processing & Management, 60:103120.

Lazer, D. M. J., Baum, M. A., Benkler, Y., Berinsky, A. J., Greenhill, K. M., Menczer, F., Metzger, M. J., Nyhan, B., Pennycook, G., Rothschild, D., et al. (2018). The science of fake news. Science, 359(6380):1094–1096.

Ren, S., Liu, Y., Zhu, Y., Bing, W., Ma, H., and Wang, W. (2024). MMSFD: Multi-grained and multi-modal fusion for short video fake news detection. In Proceedings of the 7th International Conference on Data Science and Information Technology (DSIT 2024), pages 1–11. Accessed: 2026-03-28.

Rodrigues, T. M., Bonone, L., and Mielli, R. (2020). Desinformação e crise da democracia no brasil: é possível regular fake news. Confluências— Revista Interdisciplinar de Sociologia e Direito, 22(3):30–52.

Shang, L., Kou, Z., Zhang, Y., and Wang, D. (2021). A multimodal misinformation detector for COVID-19 short videos on tiktok. In Proceedings of the 2021 IEEE International Conference on Big Data (Big Data), pages 899–908. Accessed: 2026-03-28.

Shu, K., Sliva, A., Wang, S., Tang, J., and Liu, H. (2017). Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter, 19(1):22–36.

Wardle, C. and Derakhshan, H. (2017). Information disorder: Toward an interdisciplinary framework for research and policymaking, volume 27. Council of Europe Strasbourg.

Wu, P., Zhan, Y., Zhang, L., Wang, Z., and Xu, Z. (2021). Multimodal fusion with co-attention networks for fake news detection. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2560–2569. Association for Computational Linguistics.
Publicado
08/09/2026
MORI, Adriel L. V.; LIMA, Eliomar Araújo; GRACIANO-NETO, Valdemar V.; BARBOSA, Jacson R.; CARVALHO, Rogerio Rodrigues; GALVÃO FILHO, Arlindo R.. Comparative Analysis of Language Models in the Understanding of Short Videos: A Multimodal Approach for Disinformation Detection. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 345-358. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249221.