Heurísticas vs. LLMs na triagem de tweets: avaliando pistas linguísticas para tarefas de PLN em português
Resumo
Este artigo apresenta um pipeline de triagem inteligente para pré-processamento de textos informais em tarefas de Processamento de Linguagem Natural (PLN). O pipeline pontua a utilidade textual usando pistas linguísticas fundamentadas na dimensão retórica da argumentação (clareza, credibilidade, apelo emocional e organização). Uma avaliação intrínseca comparou o pipeline proposto com um Large Language Model (LLM) utilizando um padrão ouro humano. Os resultados mostram que o pipeline obteve maior precisão em características objetivas, enquanto elementos subjetivos exigem modelos contextuais. Esses resultados indicam que uma arquitetura híbrida pode ser uma estratégia mais eficaz para a curadoria de dados textuais.
Referências
Albalak, A., Elazar, Y., Xie, S. M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., Raffel, C., Chang, S., Hashimoto, T., and Wang, W. Y. (2024). A survey on data selection for language models.
Bowman, S. R. (2024). Eight Things to Know about Large Language Models. Critical AI, 2(2).
Caseli, H. d. M., Nunes, M. d. G. V., and Pagano, A. (2024). O que é PLN? In Caseli, H. M. and Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, chapter 1. BPLN, 3 edition.
Chen, H. (2022). Data Quality Evaluation and Improvement for Machine Learning. PhD thesis, University of North Texas, USA.
Hartmann, N., Fonseca, E., Shulby, C., Treviso, M., Rodrigues, J., and Aluisio, S. (2017). Portuguese Word Embeddings: Evaluating on Word Analogies and Natural Language Tasks. In Proceedings of the 11th Brazilian Symposium in Information and Human Language Technology, Uberlândia, Brazil.
Iasulaitis, S., Valejo, A. D. B., Greco, B. C., Perillo, V. G., Messias, G. H., Vicari, I., and Interfaces—Center for Sociopolitical Studies of Algorithms and Artificial Intelligence (2025). The Interfaces Twitter Elections Dataset: Construction process and characteristics of big social data during the 2022 presidential elections in Brazil. PLOS ONE, 20(2):e0316626.
Leal, S. E., Duran, M. S., Scarton, C. E., Hartmann, N. S., and Aluísio, S. M. (2024). NILC-Metrix: Assessing the Complexity of Written and Spoken Language in Brazilian Portuguese. Language Resources and Evaluation, 58(1):73–110.
Leite, J. A., Silva, D. F., Bontcheva, K., and Scarton, C. (2020). Toxic Language Detection in Social Media for Brazilian Portuguese: New Dataset and Multilingual Analysis. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Suzhou, China.
Lopes, L., Duran, M. S., Fernandes, P., and Pardo, T. A. S. (2022). PortiLexicon-UD: A Portuguese Lexical Resource According to Universal Dependencies Model. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6635–6643, Marseille, France. European Language Resources Association.
Pinto, A., Gonçalo Oliveira, H., Figueira, Á., and Alves, A. O. (2017). Predicting the relevance of social media posts based on linguistic features and journalistic criteria. New Generation Computing, 35(4):451–472.
Silva, C. F. d. (2023). Abordagem computacional para avaliação automática da qualidade da argumentação na dimensão retórica de tweets no domínio da política brasileira. PhD thesis, Universidade Federal de São Carlos, São Carlos.
Wachsmuth, H., Naderi, N., Hou, Y., Bilu, Y., Prabhakaran, V., Thijm, T. A., Hirst, G., and Stein, B. (2017). Computational Argumentation Quality Assessment in Natural Language. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 176–187, Valencia, Spain. Association for Computational Linguistics.
