Classificação de Textos com Matrizes de Similaridade: Um Estudo de Caso em um Corpus com Nomes de Produtos
Resumo
Este trabalho aborda o problema da comparação semântica de pares de strings curtas em português. Mais especificamente, o trabalho avalia uma abordagem não supervisionada baseada em matrizes de similaridade para classificar nomes de produtos contendo diferentes tipos de especificações. As matrizes foram utilizadas de forma isolada e combinadas aos pares. Em experimentos realizados sobre um corpus composto por 1.000 instâncias, a melhor acurácia foi obtida através da combinação da matriz TF-IDF com matrizes de embeddings pré-treinados da família BERT.
Referências
Caseli, H. M., Nunes, M. G. V., and Pagano, A. (2026). O que é pln? In Caseli, H. M. and Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português – Volume 1, book chapter 1. BPLN, 4 ed.
Cer, D., et al. (2018). Universal sequence encoder. arXiv preprint arXiv:1803.11175. Disponível em: [link].
de Lima, L. S. G. and Gonçalves, E. C. (2022). Similaridade Semântica de Nomes de Produtos Alimentícios Utilizando Wordnets do Português. In Proceedings of the 15th Seminar on Ontology Research in Brazil (ONTOBRAS) and 6th Doctoral and Masters Consortium on Ontologies (WTDO), Online, CEUR-WS.org.
Google (2020). BERT. Disponível em: [link]
Hartmann, N. S., et al. (2017). Portuguese Word Embeddings: Evaluating on Word Analogies and Natural Language Tasks In Proceedings of the 9th Brazilian Symposium in Information and Human Language Technology (STIL 2017), pages 122–131, Uberlândia, SBC.
Honnibal, M. and Montani, I. (2017). “spaCy: Industrial-strength Natural Language Processing in Python”, [link].
Huang, J.-T., Sharma, A., Sun, S., Xia, X., Zhang, D., et al. (2020). Embedding-based Retrieval in Facebook Search. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 2553–256, Online. ACM.
Hugging Face. (2026). “Models”, [link]
IBGE (2016). Para compreender o INPC (um texto simplificado), 7a. ed, IBGE.
IBGE (2026). “Inflação”, [link].
Jurafsky, D. and Martin, J. H. (2026) “Embeddings”. In: Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, with Language Models. 3rd ed. Stanford University, 2026. Cap. 5. Disponível em: [link]. Acesso em: 14 mai. 2026.
Köpcke, H. and Rahm, E. (2010). Frameworks for entity matching: A comparison. Data & Knowledge Engineering, 69(2): 197–210. Elsevier
Manning, C. D., Raghavan, P., and Schütze, H. (2008). Introduction to information retrieval, Cambridge University Press.
Meirelles, T. P., Gonçalves, E. C., and Gomes, D. T. (2022). Pareamento de Nomes de Produtos e Serviços Utilizando Medidas de Similaridade Textual nos Níveis Alfabético, Léxico e Semântico. Cadernos Do IME - Série Informática, 46:104–117. IME / UERJ.
Pedregosa, F., et al. (2011). Scikit-learn: Machine learning in python. Journal of Machine Learning Research, 12:2825–2830. MIT Press.
Pellegrini, L., Santos, F., and Cantarino, F. (2025). Classificação de Notícias Falsas na Língua Portuguesa Utilizando Modelos baseados na Arquitetura Transformer. In Proceedings of the 16th Brazilian Symposium in Information and Human Language Technology (STIL 2025), pages 549-556, Fortaleza. SBC.
Reimers, N. and Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084. Disponível em: [link].
Romualdo, A., Real, L., and Caseli, H. (2021). Measuring Brazilian Portuguese Product Titles Similarity Using Embeddings. In Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology (STIL 2021), pages 121-132, Online. ACL.
Souza, F., Nogueira, R. and Lotufo, R. (2020). BERTimbau: Pretrained BERT Models for Brazilian Portuguese. In Proceedings of the 2020 Brazilian Conference on Intelligent Systems (BRACIS 2020), pages 403–417, Online. Springer.
Winkler, W. E. (1990). String Comparator Metrics and Enhanced Decision Rules in the Fellegi-Sunter Model of Record Linkage”. In Proceedings of the Sect. on Surv. Research, pages 354–359, ERIC.
Yepez, J., Tavares, B., Peres, F., and Becker, K. (2024). Na Batida do Funk: Modelagem de Tópicos Combinando LLM, Engenharia de Prompt e BERTopic. In Anais do XXXIX Simpósio Brasileiro de Bancos de Dados (SBBD 2024), pages 613–625, Florianópolis, SBC.
