Legal Nugget Extraction for Granular Retrieval over Long Jurisprudential Texts
Resumo
Legal retrieval over jurisprudential collections is challenging because court decisions are long, heterogeneous documents whose relevant legal thesis may occupy only a small portion of the text. This paper asks whether legal nuggets, defined as short and self-contained legal theses extracted from source documents, can improve dense retrieval over Brazilian legal collections. We propose a pipeline that extracts nuggets from each document, indexes them with embeddings, retrieves nugget-level evidence, and aggregates the retrieved nuggets back to document-level rankings. We evaluate this approach on four Portuguese legal retrieval benchmarks from the JUA ecosystem, reporting NDCG@10, MAP@10, and MRR@10. Nugget retrieval substantially improves the two jurisprudential datasets: on JUA-Juris, NDCG@10 increases from 0.10265 to 0.20461, and on JurisTCU from 0.20898 to 0.32696. However, it underperforms full-document retrieval on NormasTCU and BR-TaxQA, and an embedding-model ablation shows that strong domain-adapted retrievers can remain better in the full-document setting. The results demonstrate that legal nuggets can be useful for jurisprudence search, especially when queries are formulated as legal theses, but they may not transfer equally well to other legal retrieval scenarios.Referências
Arabzadeh, N. and Clarke, C. L. A. (2025). Benchmarking LLM-based relevance judgment methods. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 3194–3204.
Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., and Androutsopoulos, I. (2020). LEGAL-BERT: The muppets straight out of law school. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2898–2904.
Dietz, L., Li, B., Liu, G., Ju, J.-H., Yang, E., Lawrie, D., Walden, W., and Mayfield, J. (2026). Incorporating Q&A nuggets into retrieval-augmented generation.
Farzi, N., Dietz, L., and Lewis, D. D. (2026). Supporting humans in evaluating AI summaries of legal depositions. In Proceedings of the 2026 Conference on Human Information Interaction and Retrieval, pages 350–354.
Fernandes, L. C., Ribeiro, L. d. S., de Castro, M. V. B., da Silva Pacheco, L. A., and de Oliveira Sandes, E. F. (2026). JurisTCU: a Brazilian Portuguese information retrieval dataset with query relevance judgments. Language Resources and Evaluation, 60(1):23.
Järvelin, K. and Kekäläinen, J. (2002). Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4):422–446.
Júnior, J. D., Faria, A., de Oliveira, E. S., de Brito, E., Teotonio, M., Assumpção, A., Carmo, D., Lotufo, R., and Pereira, J. (2026). BR-TaxQA-R: A dataset for question answering with references for brazilian personal income tax law, including case law. In de Freitas, R. and Furtado, D., editors, Intelligent Systems, pages 208–222, Cham. Springer Nature Switzerland.
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 6769–6781.
Khattab, O. and Zaharia, M. (2020). ColBERT: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 39–48.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Kuttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474.
Ma, Y., Wu, Y., Ai, Q., Liu, Y., Shao, Y., Zhang, M., and Ma, S. (2024). Incorporating structural information into legal case retrieval. ACM Transactions on Information Systems, 42(2):1–28.
Manning, C. D., Raghavan, P., and Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press, Cambridge.
Pereira, J., Fernandes, L., de Brito, E., Lotufo, R., and Bonifacio, L. (2026a). JUÁ - a benchmark for information retrieval in brazilian legal text collections.
Pereira, J., Lotufo, R., and Bonifacio, L. (2026b). Domain-adaptive dense retrieval for brazilian legal search.
Thakur, N., Lin, J., Havens, S., Carbin, M., Khattab, O., and Drozdov, A. (2025). Fresh-Stack: Building realistic benchmarks for evaluating retrieval on technical documents. In Advances in Neural Information Processing Systems.
Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., and Gurevych, I. (2021). BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.
Voorhees, E. M. and Tice, D. M. (2000). Building a question answering test collection. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 200–207.
Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., and Zhou, J. (2025). Qwen3 Embedding: Advancing text embedding and reranking through foundation models.
Łajewska, W. and Balog, K. (2025). GINGER: Grounded information nugget-based generation of responses. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2723–2727.
Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., and Androutsopoulos, I. (2020). LEGAL-BERT: The muppets straight out of law school. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2898–2904.
Dietz, L., Li, B., Liu, G., Ju, J.-H., Yang, E., Lawrie, D., Walden, W., and Mayfield, J. (2026). Incorporating Q&A nuggets into retrieval-augmented generation.
Farzi, N., Dietz, L., and Lewis, D. D. (2026). Supporting humans in evaluating AI summaries of legal depositions. In Proceedings of the 2026 Conference on Human Information Interaction and Retrieval, pages 350–354.
Fernandes, L. C., Ribeiro, L. d. S., de Castro, M. V. B., da Silva Pacheco, L. A., and de Oliveira Sandes, E. F. (2026). JurisTCU: a Brazilian Portuguese information retrieval dataset with query relevance judgments. Language Resources and Evaluation, 60(1):23.
Järvelin, K. and Kekäläinen, J. (2002). Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4):422–446.
Júnior, J. D., Faria, A., de Oliveira, E. S., de Brito, E., Teotonio, M., Assumpção, A., Carmo, D., Lotufo, R., and Pereira, J. (2026). BR-TaxQA-R: A dataset for question answering with references for brazilian personal income tax law, including case law. In de Freitas, R. and Furtado, D., editors, Intelligent Systems, pages 208–222, Cham. Springer Nature Switzerland.
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 6769–6781.
Khattab, O. and Zaharia, M. (2020). ColBERT: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 39–48.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Kuttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474.
Ma, Y., Wu, Y., Ai, Q., Liu, Y., Shao, Y., Zhang, M., and Ma, S. (2024). Incorporating structural information into legal case retrieval. ACM Transactions on Information Systems, 42(2):1–28.
Manning, C. D., Raghavan, P., and Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press, Cambridge.
Pereira, J., Fernandes, L., de Brito, E., Lotufo, R., and Bonifacio, L. (2026a). JUÁ - a benchmark for information retrieval in brazilian legal text collections.
Pereira, J., Lotufo, R., and Bonifacio, L. (2026b). Domain-adaptive dense retrieval for brazilian legal search.
Thakur, N., Lin, J., Havens, S., Carbin, M., Khattab, O., and Drozdov, A. (2025). Fresh-Stack: Building realistic benchmarks for evaluating retrieval on technical documents. In Advances in Neural Information Processing Systems.
Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., and Gurevych, I. (2021). BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.
Voorhees, E. M. and Tice, D. M. (2000). Building a question answering test collection. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 200–207.
Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., and Zhou, J. (2025). Qwen3 Embedding: Advancing text embedding and reranking through foundation models.
Łajewska, W. and Balog, K. (2025). GINGER: Grounded information nugget-based generation of responses. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2723–2727.
Publicado
19/10/2026
Como Citar
PEREIRA, Lucas; BRITO, Erick; LOTUFO, Roberto; PEREIRA, Jayr.
Legal Nugget Extraction for Granular Retrieval over Long Jurisprudential Texts. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 305-316.
DOI: https://doi.org/10.5753/stil.2026.26588.
