Heurísticas vs. LLMs na triagem de tweets: avaliando pistas linguísticas para tarefas de PLN em português

  • Laura Pessine Teixeira UFSCar
  • Helena de Medeiros Caseli UFSCar

Resumo


Este artigo apresenta um pipeline de triagem inteligente para pré-processamento de textos informais em tarefas de Processamento de Linguagem Natural (PLN). O pipeline pontua a utilidade textual usando pistas linguísticas fundamentadas na dimensão retórica da argumentação (clareza, credibilidade, apelo emocional e organização). Uma avaliação intrínseca comparou o pipeline proposto com um Large Language Model (LLM) utilizando um padrão ouro humano. Os resultados mostram que o pipeline obteve maior precisão em características objetivas, enquanto elementos subjetivos exigem modelos contextuais. Esses resultados indicam que uma arquitetura híbrida pode ser uma estratégia mais eficaz para a curadoria de dados textuais.

Referências

Al Sharou, K., Li, Z., and Specia, L. (2021). Towards a better understanding of noise in natural language processing. In Mitkov, R. and Angelova, G., editors, Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), pages 53–62, Held Online. INCOMA Ltd.

Albalak, A., Elazar, Y., Xie, S. M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., Raffel, C., Chang, S., Hashimoto, T., and Wang, W. Y. (2024). A survey on data selection for language models.

Bowman, S. R. (2024). Eight Things to Know about Large Language Models. Critical AI, 2(2).

Caseli, H. d. M., Nunes, M. d. G. V., and Pagano, A. (2024). O que é PLN? In Caseli, H. M. and Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, chapter 1. BPLN, 3 edition.

Chen, H. (2022). Data Quality Evaluation and Improvement for Machine Learning. PhD thesis, University of North Texas, USA.

Hartmann, N., Fonseca, E., Shulby, C., Treviso, M., Rodrigues, J., and Aluisio, S. (2017). Portuguese Word Embeddings: Evaluating on Word Analogies and Natural Language Tasks. In Proceedings of the 11th Brazilian Symposium in Information and Human Language Technology, Uberlândia, Brazil.

Iasulaitis, S., Valejo, A. D. B., Greco, B. C., Perillo, V. G., Messias, G. H., Vicari, I., and Interfaces—Center for Sociopolitical Studies of Algorithms and Artificial Intelligence (2025). The Interfaces Twitter Elections Dataset: Construction process and characteristics of big social data during the 2022 presidential elections in Brazil. PLOS ONE, 20(2):e0316626.

Leal, S. E., Duran, M. S., Scarton, C. E., Hartmann, N. S., and Aluísio, S. M. (2024). NILC-Metrix: Assessing the Complexity of Written and Spoken Language in Brazilian Portuguese. Language Resources and Evaluation, 58(1):73–110.

Leite, J. A., Silva, D. F., Bontcheva, K., and Scarton, C. (2020). Toxic Language Detection in Social Media for Brazilian Portuguese: New Dataset and Multilingual Analysis. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Suzhou, China.

Lopes, L., Duran, M. S., Fernandes, P., and Pardo, T. A. S. (2022). PortiLexicon-UD: A Portuguese Lexical Resource According to Universal Dependencies Model. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6635–6643, Marseille, France. European Language Resources Association.

Pinto, A., Gonçalo Oliveira, H., Figueira, Á., and Alves, A. O. (2017). Predicting the relevance of social media posts based on linguistic features and journalistic criteria. New Generation Computing, 35(4):451–472.

Silva, C. F. d. (2023). Abordagem computacional para avaliação automática da qualidade da argumentação na dimensão retórica de tweets no domínio da política brasileira. PhD thesis, Universidade Federal de São Carlos, São Carlos.

Wachsmuth, H., Naderi, N., Hou, Y., Bilu, Y., Prabhakaran, V., Thijm, T. A., Hirst, G., and Stein, B. (2017). Computational Argumentation Quality Assessment in Natural Language. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 176–187, Valencia, Spain. Association for Computational Linguistics.
Publicado
19/10/2026
TEIXEIRA, Laura Pessine; CASELI, Helena de Medeiros. Heurísticas vs. LLMs na triagem de tweets: avaliando pistas linguísticas para tarefas de PLN em português. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 626-631. DOI: https://doi.org/10.5753/stil.2026.29820.