Balancing Coherence and Granularity in Legal Topic Modeling: A BERTopic-based Study on Portuguese Appellate Court Decisions
Resumo
The growing backlog of the Brazilian Judiciary demands scalable semantic organization beyond keyword-based retrieval. This study investigates the trade-off between topic coherence and thematic granularity in neural topic modeling for legal texts. We cluster 1,831 Brazilian labor summaries using BERTopic with LegalBERT-pt embeddings and evaluate the impact of the min_topic_size parameter on coherence (Cv), granularity, and diversity. We propose a balanced evaluation score integrating coherence and granularity. Results show higher coherence (Cv ≈ 0.78) than Latent Dirichlet Allocation (Cv ≈ 0.49), while enabling controlled topic resolution. These findings provide a practical criterion for tuning topic models in large-scale legal systems.
Referências
Blei, D. M., Ng, A. Y. and Jordan, M. I. (2003) “Latent Dirichlet Allocation”, Journal of Machine Learning Research, vol. 3, pp. 993–1022. Available at: [link].
Borges, B. R. (2025) “Comparison of Clustering Techniques in Text Documents in Portuguese”, ISys - Brazilian Journal of Information Systems, vol. 18, no. 1, pp. 4:1–4:17. DOI: 10.5753/isys.2025.5029.
Bouma, G. (2009) “Normalized (Pointwise) Mutual Information in Collocation Extraction”. Available at: [link].
Campagnolo, J. M., Duarte, D. and Dal Bianco, G. (2022) “Topic coherence metrics: How sensitive are they?”, Journal of Information and Data Management, vol. 13, no. 4, pp. 476–491. DOI: 10.5753/jidm.2022.2181.
Campello, R., Moulavi, D. and Sander, J. (2013) “Density-Based Clustering Based on Hierarchical Density Estimates”, In: Advances in Knowledge Discovery and Data Mining, Springer. DOI: 10.1007/978-3-642-37456-2_14.
Conselho Nacional de Justiça (CNJ) (2025) Justiça em Números 2025: Ano-base 2024. Departamento de Pesquisas Judiciárias, Brasília, DF.
da Silva, L., Rodrigues, M., Archanjo, A., Pessoa, L., Silva, M., de Almeida, T., & Silveira, L. (2024). Segmentação Textual Baseada em Tópicos em Português Utilizando BERTimbau. In Anais do XV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, (pp. 32-36). Porto Alegre: SBC. DOI: 10.5753/stil.2024.245080.
Dieng, A. B., Ruiz, F. J. R. and Blei, D. M. (2020) “Topic Modeling in Embedding Spaces”, Transactions of the Association for Computational Linguistics, vol. 8, pp. 439–453. DOI: 10.1162/tacl_a_00325.
Dragnut, E. et al. (2009) “Stop word and related problems in web interface integration”, Proceedings of the VLDB Endowment. DOI: 10.14778/1687627.1687667.
Glez-Peña, D. et al. (2014) “Web scraping technologies in an API world”. Briefings in Bioinformatics. DOI: 10.1093/bib/bbt026.
Grootendorst, M. (2022) “BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure”, arXiv preprint. DOI: 10.48550/arXiv.2203.05794.
Hussain, M., Rehman, U. U., Nguyen, T. D. T. and Lee, S. (2023) “CoT-STS: A Zero-shot Chain-of-thought Prompting for Semantic Textual Similarity”, In: Proceedings of the 6th Artificial Intelligence and Cloud Computing Conference (AICCC), Kyoto, Japan. ACM. DOI: 10.1145/3639592.3639611.
Joshi, A., Kale, S., Chandel, S., & Pal, D. K. (2015). Likert Scale: Explored and Explained. Current Journal of Applied Science and Technology, 7(4), 396–403. DOI: 10.9734/BJAST/2015/14975.
Martins, V. S. and Silva, C. D. (2021) “Text Classification in Law Area: A Systematic Review”, In: Proceedings of the Symposium on Knowledge Discovery, Mining and Learning (KDMiLe). DOI: 10.5753/kdmile.2021.17458.
Mimno, D. et al. (2009) “Optimizing Semantic Coherence in Topic Models”, In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). [link].
Mitchell, R. (2024) Web Scraping com Python (3rd ed.). São Paulo: Novatec Editora.
Novaes, L. P., Vianna, D. and da Silva, A. S. (2023) “Modelagem de Tópicos para a Tarefa de Recuperação de Casos Legais”, In: Proceedings of the 38th Simpósio Brasileiro de Banco de Dados (SBBD), Belo Horizonte, Brazil. Sociedade Brasileira de Computação. DOI: 10.5753/sbbd.2023.232576.
Röder, M., Both, A. and Hinneburg, A. (2015) “Exploring the Space of Topic Coherence Measures”, In: Proceedings of the 8th ACM International Conference on Web Search and Data Mining (WSDM), Shanghai, China. ACM, pp. 399–408. Available at: [link].
Sarica, S. and Luo, J. (2021) “Stopwords in Technical Language Processing”, PLOS ONE, vol. 16, no. 7. DOI: 10.1371/journal.pone.0254937.
Silva, M. et al. (2024) “Evaluating Domain-adapted Language Models for Governmental Text Classification Tasks in Portuguese”, In: Proceedings of the 39th Brazilian Symposium on Databases (SBBD), Florianópolis, Brazil. DOI: 10.5753/sbbd.2024.240508.
Souza, F., Nogueira, R., and Lotufo, R. (2020). Bertimbau: pretrained bert models for brazilian portuguese. In Proceedings of the Brazilian Conference on Intelligent Systems, pages 403–417. Springer. DOI: 10.1007/978-3-030-61377-8_28.
Silveira, R., Ponte, C., Almeida, V., Pinheiro, V. and Furtado, V. (2023) “LegalBERT-pt: A Pretrained Language Model for the Brazilian Portuguese Legal Domain”, In: Proceedings of the 12th Brazilian Conference on Intelligent Systems (BRACIS), Belo Horizonte, Brazil. Sociedade Brasileira de Computação. Available at: [link].
