Um Pipeline Data-Centric para Construção de Corpus de Reconhecimento de Entidades Nomeadas para o Domínio Trabalhista Brasileiro

  • Emerson Diego da Costa Araújo Instituto Federal da Paraíba (IFPB) / Tribunal Regional do Trabalho da 13ª Região (TRT-13)
  • Diego Ernesto Rosa Pessoa Instituto Federal da Paraíba (IFPB)
  • Hildeberg Oliveira Albuquerque Universidade Federal Rural de Pernambuco (UFRPE)

Resumo


Este trabalho apresenta um pipeline data-centric para construção e curadoria de corpora de Reconhecimento de Entidades Nomeadas na Justiça do Trabalho brasileira. A abordagem integra geração sintética, supervisão fraca, aprendizado ativo e validação humana, resultando em dois recursos: um corpus de treinamento com 53.324 segmentos anotados em 32 classes e um conjunto Gold Standard independente. A qualidade do corpus foi validada por ajuste fino do BERTimbau, que atingiu 98,15% de micro-F1 no Gold Standard. Em comparação com Qwen-2.5-72B, a solução especializada reduziu a latência de inferência em aproximadamente 1.500× (31 s para 0,02 s por segmento), evidenciando viabilidade operacional em infraestrutura institucional local.
Palavras-chave: Reconhecimento de Entidades Nomeadas, Data-Centric AI, Justiça do Trabalho, Processamento de Linguagem Natural, BERTimbau

Referências

Albuquerque, H. O., Costa, R., Silvestre, G., Souza, E., da Silva, N. F. F., Vitório, D., Moriyama, G., Martins, L., Soezima, L., Nunes, A., Siqueira, F., Tarrega, J. P., Beinotti, J. V., Dias, M., Silva, M., Gardini, M., Silva, V., de Carvalho, A. C. P. L. F., and Oliveira, A. L. I. (2022). UlyssesNER-Br: A corpus of Brazilian legislative documents for named entity recognition. In Computational Processing of the Portuguese Language (PROPOR 2022), pages 3–14, Fortaleza, Brazil. Springer.

Albuquerque, H. O., Souza, E., Lucena, D. C. G., Albuquerque, H. J. O., Silva, N. F. F. d., Dias, M. d. S., Nunes, R. O., Oliveira, A. L. I., and Carvalho, A. C. P. L. F. d. (2026). UlyssesLegalNER-Br: from legislative to legal, a comprehensive corpus of Brazilian legal documents for named entity recognition. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 331–341, Salvador, Brazil. Association for Computational Linguistics.

Andrade, C. M. V. d., França, C., Belém, F., Jallais, G., Ganem, M. A. S., Teixeira, G., Laender, A. H. F., and Gonçalves, M. A. (2023). PromptNER: Uma abordagem para reconhecimento de entidades nomeadas em dados sensíveis a partir de instâncias rotuladas automaticamente. In Anais do XXXVIII Simpósio Brasileiro de Bancos de Dados (SBBD 2023), pages 269–281, Belo Horizonte, MG, Brasil. SBC.

Araujo, P. H. L. d., Campos, T. E. d., Oliveira, R. R. R. d., Stauffer, M., Couto, S., and Bermejo, P. (2018). LeNER-Br: a dataset for named entity recognition in Brazilian legal text. In Computational Processing of the Portuguese Language (PROPOR 2018), pages 313–323. Springer.

Ariai, F., Mackenzie, J., and Demartini, G. (2025). Natural language processing for the legal domain: A survey of tasks, datasets, models, and challenges. ACM Computing Surveys, 58(6):1–37.

Brasil (2018). Lei nº 13.709, de 14 de agosto de 2018 (Lei Geral de Proteção de Dados Pessoais).

Brito, A. M., Pinheiro, V., Furtado, V., Monteiro Neto, J. A., Bomfim, F. d. C. J., da Costa, A. C. F., Silveira, R., and Aragão, N. (2023). CDJUR-BR — Uma Coleção Dourada do Judiciário Brasileiro com Entidades Nomeadas Refinadas. In Proceedings of the 14th Brazilian Symposium in Information and Human Language Technology (STIL 2023), pages 176–195, Belo Horizonte, Brazil. Association for Computational Linguistics.

Castro, P. V. Q. d. (2019). Aprendizagem profunda para reconhecimento de entidades nomeadas em domínio jurídico. Master’s thesis, Universidade Federal de Goiás, Goiânia.

Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1):37–46.

Dias, M., Boné, J., Ferreira, J. C., Ribeiro, R., and Maia, R. (2020). Named entity recognition for sensitive data discovery in Portuguese. Applied Sciences, 10(7):2303.

Ding, B., Qin, C., Zhao, R., Luo, T., Li, X., Chen, G., Xia, W., Hu, J., Luu, A. T., and Joty, S. (2024). Data augmentation using LLMs: Data perspectives, learning paradigms and challenges. In Findings of the Association for Computational Linguistics: ACL 2024, pages 1679–1705, Bangkok, Thailand. Association for Computational Linguistics.

Jurafsky, D. and Martin, J. H. (2024). Speech and Language Processing. Pearson, 3rd ed. draft edition.

Maffeo, G., Silva, C., and Oliveira, H. G. (2026). Prompt engineering for named entity extraction from Portuguese legal documents. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 1092–1097, Salvador, Brazil. Association for Computational Linguistics.

Mosbach, M., Pimentel, T., Ravfogel, S., Klakow, D., and Cotterell, R. (2023). Few-shot fine-tuning vs. in-context learning: A fair comparison and evaluation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 12284–12314. Association for Computational Linguistics.

Nunes, R. O. (2025). Data contamination in specialized named entity recognition corpora. Master’s thesis, Universidade Federal do Rio Grande do Sul, Porto Alegre.

Nunes, R. O., Balreira, D. G., Spritzer, A. S., and Freitas, C. M. D. S. (2024). A named entity recognition approach for Portuguese legislative texts using self-learning. In Proceedings of the 16th International Conference on Computational Processing of Portuguese (PROPOR 2024), pages 290–300. Springer.

Souza, F., Nogueira, R., and Lotufo, R. (2020). BERTimbau: Pretrained BERT models for Brazilian Portuguese. In Intelligent Systems: 9th Brazilian Conference (BRACIS 2020), pages 403–417. Springer.

Tran, H. T. H., Chatterjee, N., Pollak, S., and Doucet, A. (2024). DeBERTa beats behemoths: A comparative analysis of fine-tuning, prompting, and PEFT approaches on LegalLensNER. In Proceedings of the Natural Legal Language Processing Workshop 2024, pages 371–380, Miami, FL, USA. Association for Computational Linguistics.

Vithanage, D., Yu, P., Xie, Q., Xu, H., Wang, L., and Deng, C. (2025). A comprehensive evaluation of large language models for information extraction from unstructured electronic health records in residential aged care. Computers in Biology and Medicine, 197:111013.

Zha, D., Bhat, Z. P., Lai, K.-H., Yang, F., Jiang, Z., Zhong, S., and Hu, X. (2025). Data-centric artificial intelligence: A survey. ACM Computing Surveys, 57(5).

Zhong, H., Xiao, C., Tu, C., Zhang, T., Liu, Z., and Sun, M. (2020). How does NLP benefit legal system: A summary of legal artificial intelligence. arXiv preprint arXiv:2004.12158.
Publicado
08/09/2026
ARAÚJO, Emerson Diego da Costa; PESSOA, Diego Ernesto Rosa; ALBUQUERQUE, Hildeberg Oliveira. Um Pipeline Data-Centric para Construção de Corpus de Reconhecimento de Entidades Nomeadas para o Domínio Trabalhista Brasileiro. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 467-480. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249240.