Extração de Esquemas em Documentos Legais Não Estruturados Utilizando LLMs

  • Luan Alecxander Krzyzaniak UFFS
  • Denio Duarte UFFS
  • Geomar A. Schreiner UFFS
  • Giancarlo Salton UFFS

Resumo


Este trabalho explora o uso de Large Language Models (LLMs) na extração de esquemas a partir de documentos legais (atas de registro de preços) não estruturados. O estudo propõe um pipeline para aplicação de LLMs na identificação e organização eficiente de informações jurídicas. Para avaliação, foi utilizado o modelo Mistral 7B Instruct v0.2 em 471 documenos, segmentadas em blocos de 2.000 tokens. Os resultados da análise por documento indicam alto desempenho na classificação de tipos de dados (Type Accuracy: 0,946; Type Precision: 0,972), mas desempenho semântico moderado (Semantic Accuracy: 0,412; Semantic Coverage: 0,714), revelando consistência na tipagem, porém limitações na correspondência semântica.

Referências

Aquino, I., dos Santos, M. M., Dorneles, C., and Carvalho, J. T. (2024). Extracting information from brazilian legal documents with retrieval augmented generation. In Anais Estendidos do XXXIX SBBD, pages 280–287, Porto Alegre, RS, Brasil. SBC.

Banhara, N., Schreiner, G. A., da Silva Feitosa, S., and Duarte, D. (2024). Enumeration, tagged unions, tuples, and collections: A novel approach to extracting json schema. In Simpósio Brasileiro de Banco de Dados (SBBD), pages 234–246. SBC.

Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of ACM FAccT, pages 610–623.

Christopher, C., Moore, K., and Liebowitz, D. (2022). SchemaDB: A Dataset for Structures in Relational Data, page 233–243. Springer Nature Singapore.

Dagdelen, J. et al. (2024). Structured information extraction from scientific text with large language models. Nature communications, 15(1):1418.

Jiang, A. Q. et al. (2023). Mistral 7b.

Lewis, P. et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In 34th International Conference on Neural Information Processing Systems, NIPS ’20.

Lima, E. d. S. and Flores, D. (2016). A evolução da legislação relacionada à digitalização e aos documentos digitais no âmbito da administração pública federal. Revista Sociais e Humanas, 29(1):75–91.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.

Wei, J. et al. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.

Zhang, B. and Soh, H. (2024). Extract, define, canonicalize: An LLM-based framework for knowledge graph construction. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.
Publicado
22/04/2026
KRZYZANIAK, Luan Alecxander; DUARTE, Denio; SCHREINER, Geomar A.; SALTON, Giancarlo. Extração de Esquemas em Documentos Legais Não Estruturados Utilizando LLMs. In: ESCOLA REGIONAL DE BANCO DE DADOS (ERBD), 21. , 2026, Dois Vizinhos/PR. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 49-58. ISSN 2595-413X. DOI: https://doi.org/10.5753/erbd.2026.21041.