Além das Palavras: Detectando Técnicas Persuasivas em Notícias e Artigos de Opinião da Língua Portuguesa com BERT-Tiny
Resumo
O presente trabalho propõe uma abordagem para a identificação automática de técnicas de persuasão em textos da língua portuguesa. Para isso, foi construída uma base de dados rotulada com Large Language Models (LLMs) em oito categorias, sendo sete técnicas de persuasão derivadas da SemEval23-T3 e uma classe neutra. A partir da base, foi realizado o fine-tuning do modelo BERT-Tiny, visando estabelecer uma baseline compacta para a tarefa de classificação. Os resultados preliminares alcançaram 86% de acurácia em ambiente de hardware limitado (GPU NVIDIA T4), indicando uma alternativa viável para o desenvolvimento de recursos de análise de estratégias persuasivas em desinformação para a língua portuguesa.
Referências
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C., and Solorio, T., editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
Kasner, Z., Zouhar, V., Schmidtová, P., Kartáč, I., Onderková, K., Plátek, O., Gkatzia, D., Mahamood, S., Dušek, O., and Balloccu, S. (2025). LLMs as span annotators: A comparative study of LLMs and humans. arXiv preprint arXiv:2504.08697.
Monteiro, R. A., Santos, R. L. S., Pardo, T. A. S., Almeida, T. A., Ruiz, E. E. S., and Vale, O. A. (2018). Contributions to the study of fake news in portuguese: New corpus and automatic detection results. In Proceedings of the 13th International Conference on Computational Processing of the Portuguese Language, pages 324–334.
Newman, N., Fletcher, R., Robertson, C. T., Ross Arguedas, A., and Nielsen, R. K. (2024). Reuters institute digital news report 2024. Technical report, Reuters Institute for the Study of Journalism, University of Oxford.
Oshikawa, R., Qian, J., and Wang, W. Y. (2020). A survey on natural language processing for fake news detection. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6086–6093, Marseille, France. European Language Resources Association.
Piskorski, J., Stefanovitch, N., Da San Martino, G., and Nakov, P. (2023). SemEval-2023 task 3: Detecting the category, the framing, and the persuasion techniques in online news in a multi-lingual setup. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), pages 2343–2361, Toronto, Canada. Association for Computational Linguistics.
Resende, G., Melo, P., Sousa, H., Messias, J., Vasconcelos, M., Almeida, J. M., and Benevenuto, F. (2019). (mis)information dissemination in whatsapp: Gathering, analyzing and countermeasures. In Proceedings of The Web Conference 2019, pages 818–828.
Tan, Z., Beigi, A., Wang, S., Guo, R., Bhattacharjee, A., Jiang, B., Karami, M., Li, J., Cheng, L., and Liu, H. (2024). Large language models for data annotation: A survey. arXiv preprint arXiv:2402.13446.
Wardle, C. and Derakhshan, H. (2017). Information disorder: Toward an interdisciplinary framework for research and policy making. Technical report, Council of Europe.
World Economic Forum (2024). The global risks report 2024. Technical report, World Economic Forum.
