Contradiction Detection in Brazilian Legislative Bills: A Limited Labeled Data Study
Resumo
Identifying contradictory legislative bills is essential for ensuring consistency in Brazilian legal frameworks, yet it remains a manual and hard-to-scale task. In this work, we study automatic contradiction detection under limited labeled data. We use a dataset of 106 contradictory bill pairs curated by domain specialists and investigate two strategies: (i) supervised learning with synthetic data generation and (ii) zero-shot and few-shot classification using generative large language models (LLMs). Results show that few-shot LLMs perform strongly and outperform the supervised approach based on synthetic data. Although the small dataset limits generalization, our findings highlight the challenges of synthetic data in legal domains and the potential of LLMs for low-data scenarios.
Referências
Brasil - Minas Gerais (2025). Regimento interno da Assembleia Legislativa do Estado de Minas Gerais. [link]. Accessed: 2026-05-01.
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT, pages 4171–4186. Association for Computational Linguistics.
Garcia, E., Silva, N., Siqueira, F., Gomes, J., Albuquerque, H. O., Souza, E., Lima, E., and de Carvalho, A. (2024). RoBERTaLexPT: A legal RoBERTa model pretrained with deduplication for Portuguese. In Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1, pages 374–383, Santiago de Compostela, Galicia/Spain. Association for Computational Lingustics.
Martins, M. and Medeiros, C. (2025). From anvisa leaflets to extended interoperability with global health databases: some pitfalls and success stories. In Anais do XIX Brazilian e-Science Workshop, pages 9–16, Porto Alegre, RS, Brasil. SBC.
Mistral AI (2025). Mistral-small-24b-instruct-2501. [link]. Accessed: 2026-05-01.
Navastara, D. A., Abdillah, S., Benito, D., Adillion, I. G., and Purwitasari, D. (2025). Document matching for contradiction detection in low-resource legislative texts with self-training and augmentation using transformer model. JANAPATI, 14(2):321–335.
OpenAI (2025a). Gpt-5 mini. [link]. Accessed: 2026-05-01.
OpenAI (2025b). Gpt-5 nano. [link]. Accessed: 2026-05-01.
Silva, M., Oliveira, G., Costa, L., and Pappa, G. (2024). Evaluating domain-adapted language models for governmental text classification tasks in Portuguese. In Anais do XXXIX SBBD, pages 247–259, Porto Alegre, RS, Brasil. SBC.
Silva, M. O., Oliveira, G. P., Costa, L. G. L., and Pappa, G. L. (2025). GovBERT-BR: A BERT-based language model for Brazilian Portuguese governmental data. In Intelligent Systems. BRACIS 2024, pages 19–32, Cham. Springer Nature Switzerland.
Souza, F., Nogueira, R., and Lotufo, R. (2020). BERTimbau: pretrained BERT models for Brazilian Portuguese. In 9th BRACIS, Rio Grande do Sul, Brazil, October 20-23.
Viegas, C. F. O., Costa, B. C., and Ishii, R. P. (2023). JurisBERT: A new approach that converts a classification corpus into an STS one. In Computational Science and Its Applications – ICCSA 2023, pages 349–365. Springer Nature Switzerland.
Zaiton, H., Alansary, S., Sarwat, N., and Hussein, I. (2026). Automatic detection of contradictions in legal texts: A computational linguistic approach. Procedia Computer Science, 275:503–512. 7th ACling.
