A Comparative Evaluation of Embedding Models for Retrieval and Clustering of Brazilian Legislative Amendments

  • João Robson Santos Martins Senado Federal / UnB

Resumo


The large volume of amendments submitted in the Brazilian legislative process makes semantic analysis and consolidation labor-intensive. We present a comparative evaluation of 14 distinct base representation models, including BM25L and embedding models spanning open-source multilingual, proprietary, Portuguese-adapted, and legislative-specialized architectures, for retrieving and clustering legislative amendments. Across three corpora totaling 2,462 documents, we assess retrieval, clustering, text segmentation, and thematic granularity. Results show that large-scale open-source multilingual models constitute the best-performing group overall, with Octen-Embedding-8B exhibiting the most consistent performance. gemini-embedding-2 is highly competitive in sentence-similarity mode, while preprocessed BM25L remains a strong baseline, outperforming specialized encoders. Finally, removing structural metadata improves performance, whereas removing justifications degrades global structure; thematic aggregation improves local accuracy at the expense of global consistency.

Palavras-chave: Amendments Clustering, Amendments Retrieval, Embeddings, Legislative Process

Referências

Agnoloni, T., Marchetti, C., Battistoni, R., and Briotti, G. Clustering Similar Amendments at the Italian Senate. In Proceedings of the Workshop ParlaCLARIN III at the 13th LREC. Marseille, France, pp. 39–46, 2022.

Andrade, P., Souza, E., Silva, N., and Carvalho, A. Semantic Clustering in the Context of Legislative Amendments. In Anais do XXII Encontro Nacional de Inteligência Artificial e Computacional. Fortaleza, Brazil, pp. 569–579, 2025.

Chung, I., Kerboua, I., Kardos, M., Solomatin, R., and Enevoldsen, K. Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks. In Proceedings of the Workshop on Championing Open-source Development in Machine Learning at the 42th ICML. Vancouver, Canada, 2025.

DeepMind, G. Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini. [link], 2026.

dos Santos, J. A., Souza, E., Bastos Filho, C. J. A., Albuquerque, H. O., Vitório, D., de Lucena, D. C. G., Silva, N., and de Carvalho, A. HIRS: A Hybrid Information Retrieval System for Legislative Documents. In Proceedings of the 23rd EPIA Conference on Artificial Intelligence, Part I. Viana do Castelo, Portugal, pp. 320–331, 2024.

Enevoldsen, K., Chung, I., Kerboua, I., Kardos, M., Mathur, A., Stap, D., Gala, J., Siblini, W., Krzemiński, D., Winata, G., Sturua, S., Utpala, S., Ciancone, M., Schaeffer, M., Misra, D., Dhakal, S., Rystrø m, J., Solomatin, R., Çağatan, O., Kundu, A., Bernstorff, M., Xiao, S., Sukhlecha, A., Pahwa, B., Poświata, R., GV, K. K., Ashraf, S., Auras, D., Plüster, B., Harries, J., Magne, L., Mohr, I., Zhu, D., Gisserot-Boukhlef, H., Aarsen, T., Kostkan, J., Wojtasik, K., Lee, T., Suppa, M., Zhang, C., Rocca, R., Hamdy, M., Michail, A., Yang, J., Faysse, M., Vatolin, A., Thakur, N., Dey, M., Vasani, D., Chitale, P., Tedeschi, S., Tai, N., Snegirev, A., Hendriksen, M., Günther, M., Xia, M., Shi, W., Lu, X. H., Clive, J., K, G., Anna, M., Wehrli, S., Tikhonova, M., Panchal, H., Abramov, A., Ostendorff, M., Liu, Z., Clematide, S., Miranda, L. J. V., Fenogenova, A., Song, G., Bin Safi, R., Li, W.-D., Borghini, A., Cassano, F., Hansen, L., Hooker, S., Xiao, C., Adlakha, V., Weller, O., Reddy, S., and Muennighoff, N. MMTEB: Massive Multilingual Text Embedding Benchmark. In International Conference on Learning Representations. pp. 101715–101771, 2025.

Gomes, L., Branco, A., Silva, J. a., Rodrigues, J. a., and Santos, R. Open Sentence Embeddings for Portuguese with the Serafim PT* Encoders Family. In Proceedings of the 23rd EPIA Conference on Artificial Intelligence, Part III. Viana do Castelo, Portugal, pp. 267–279, 2024.

Górski, Towards Legal Change Analysis: Clustering of Polish Civil Code Amendments. In Proceedings of the Third Workshop on Automated Semantic Analysis of Information in Legal Texts, co-located with the 17th International Conference on Artificial Intelligence and Law (ICAIL). Montreal, Canada, pp. 1–5, 2019.

Manning, C. D., Raghavan, P., and Schütze, H. Introduction to Information Retrieval. Cambridge University Press, 2008.

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research vol. 12, pp. 2825–2830, 2011.

Presidência da República. Medida Provisória nº 927, de 22 de março de 2020. [link], 2020.

Pressato, D., de Andrade, P. L. C., Junior, F. R., Siqueira, F. A., Souza, E. P. R., da Silva, N. F. F., de Souza Dias, M., and de Leon Ferreira de Carvalho, A. C. P. Natural Language Processing Application in Legislative Activity: a Case Study of Similar Amendments in the Brazilian Senate. In Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1. Santiago de Compostela, Galicia/Spain, pp. 614–619, 2024.

Reimers, N. and Gurevych, I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China, pp. 3982–3992, 2019.

Rosenberg, A. and Hirschberg, J. V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL). Prague, Czech Republic, pp. 410–420, 2007.

Sajeva, A., Iannucci, S., Marchetti, C., Merialdo, P., and Torlone, R. Clustering Amendments with Semantic Embeddings. In Proceedings of the 32nd Symposium of Advanced Database Systems. Villasimius, Italy, pp. 312–320, 2024.

Senado Federal. Relatório de Gestão 2025. [link], 2026.

Smith, R. An Overview of the Tesseract OCR Engine. In Proceedings of the Ninth International Conference on Document Analysis and Recognition - Volume 02. Curitiba, Brazil, pp. 629–633, 2007.

Souza, F., Nogueira, R., and Lotufo, R. BERTimbau: Pretrained BERT Models for Brazilian Portuguese. In Proceedings of the 9th Brazilian Conference on Intelligent Systems, Part I. Rio Grande, Brazil, pp. 403–417, 2020.

Vitório, D., Souza, E., Dos Santos, J. A., De Carvalho, A. C. P. d. L. F., Oliveira, A. L. I., and F. da Silva, N. F. BM25 x Vila Sésamo: avaliando modelos Sentence-BERT para Recuperação de Informação no cenário legislativo brasileiro. Linguamática 17 (1): 17–33, 2025.

Wilcoxon, F. Individual comparisons by ranking methods. Biometrics bulletin 1 (6): 80–83, 1945.
Publicado
19/10/2026
MARTINS, João Robson Santos. A Comparative Evaluation of Embedding Models for Retrieval and Clustering of Brazilian Legislative Amendments. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 9-16. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.32143.