EpiCorpus-BR: A Named Entity Recognition Corpus for Epidemiological Documents in Portuguese

Resumo


This paper presents EpiCorpus-BR, the first annotated corpus for Named Entity Recognition (NER) in Brazilian Portuguese epidemiological surveillance documents, covering the annual SIREVA-SUS reports (2013–2024) and complementary documents from the Brazilian Ministry of Health. The corpus comprises 8,314 textual units across 22 documents (6,358 narrative text sentences and 1,956 table rows), with 4,159 entity mentions in the silver set. The tagset consists of eight specialized categories (PATOGENO, SOROTIPO, FAIXA_ETARIA, LOCAL, PERIODO_TEMPORAL, METODO_LAB, METRICA_EPI, MANIFESTACAO_CLINICA), annotated via few-shot prompting with GPT-4o-mini and manually reviewed in a stratified sample of 280 units (180 text sentences and 100 table rows) by two independent annotators (κ = 0.76). Strict IOB2 evaluation with seqeval yields a combined micro-F1 of 0.77. METRICA_EPI is absent from narrative text but reaches F1 = 0.89 on table rows, which shows why the tabular sub-corpus must be included in the evaluation. A BERTimbau baseline fine-tuned on the silver corpus, with the gold held out, reaches micro-F1 = 0.73, showing that the corpus supports training a dedicated NER model. The corpus and code are publicly available.

Palavras-chave: Named Entity Recognition, Corpus, Epidemiological Surveillance, Brazilian Portuguese, Semi-automatic Annotation, SIREVA-SUS

Referências

Albuquerque, H. O., Souza, E., Gomes, C., Pinto, M. H. d. C., Filho, R. P. S., Costa, R., Lopes, V. T. d. M., da Silva, N. F. F., de Carvalho, A. C., and Sampaio, A. d. O. (2023). Named entity recognition: a survey for the Portuguese language. Procesamiento del Lenguaje Natural, 70:171–185.

Brito, M., Pinheiro, V., Furtado, V., Monteiro Neto, J. A., Bomfim, F. d. C. J., da Costa, A. C. F., and Silveira, R. (2023). CDJUR-BR – uma coleção dourada do judiciário brasileiro com entidades nomeadas refinadas. In Anais do XIV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL), pages 177–186, Porto Alegre. Sociedade Brasileira de Computação.

Ding, B., Qin, C., Zhao, R., Luo, T., Li, X., Chen, G., Xia, W., Hu, J., Luu, A. T., and Joty, S. (2024). Data augmentation using LLMs: Data perspectives, learning paradigms and challenges. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pages 8764–8786.

Doğan, R. I., Leaman, R., and Lu, Z. (2014). NCBI disease corpus: A resource for disease name recognition and concept normalization. Journal of Biomedical Informatics, 47:1–10.

DS4SD Team (2024). Docling technical report. arXiv preprint arXiv:2408.09869.

Freitas, C., Rabonato, R. T., and Berton, L. (2025). Enhancing epidemiological insights with RAG for SIREVA-SUS reports. In Anais do XXII Encontro Nacional de Inteligência Artificial e Computacional (ENIAC), pages 1364–1375, Fortaleza/CE, Brazil. Sociedade Brasileira de Computação.

Freitas, C., Real, L., Berton, L., and Paiva, V. d. (2026). Towards a Universal Dependencies corpus for Portuguese epidemiological reports. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 2, pages 228–237, Salvador, Brazil. Association for Computational Linguistics.

Landis, J. R. and Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1):159–174.

Moreira, H., Silva, P. F. d., Vieira, R., and Moreira, V. (2025). PetroGeoNER: um conjunto de dados refinado e unificado para NER no domínio de petróleo e gás. In Anais do XVI Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL), pages 259–271, Porto Alegre. Sociedade Brasileira de Computação.

Nadeau, D. and Sekine, S. (2007). A survey of named entity recognition and classification. Lingvisticae Investigationes, 30(1):3–26.

Nakayama, H. (2018). seqeval: A python framework for sequence labeling evaluation. [link].

Oliveira, L. E. S. e., Peters, A. C., da Silva, A. M. P., Gebeluca, C. P., Gumiel, Y. B., Cintho, L. M. M., Carvalho, D. R., Al Hasan, S., and Moro, C. M. C. (2022). SemClinBr – a multi-institutional and multi-specialty semantically annotated corpus for Portuguese clinical NLP tasks. Journal of Biomedical Semantics, 13(1).

OpenAI (2024). GPT-4o mini: Advancing cost-efficient intelligence. [link].

Piai, L., Di-Felippo, A., and Romano, N. T. (2025). Entidades nomeadas em tweets do mercado de ações: uma anotação refinada e com motivação linguística. In Anais do XVI Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL), pages 654–663, Porto Alegre. Sociedade Brasileira de Computação.

Pontes, M. F., Pedrosa, R. C., Lopes, P. H., and Luz, E. J. (2024). Avaliando aprendizagem federada com criptografia homomórfica para reconhecimento de entidades médicas nomeadas utilizando modelos compactos de BERT. In Anais do XV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL), pages 48–56, Porto Alegre. Sociedade Brasileira de Computação.

Ramos, F. d. O., Marvila, C. G., and Colombo, C. d. S. (2026). Uma abordagem híbrida para reconhecimento de entidades nomeadas em relatos de casos clínicos. In Anais do XXII Simpósio Brasileiro de Sistemas de Informação (SBSI), pages 216–233, Porto Alegre. Sociedade Brasileira de Computação.

Schneider, E. T. R., de Souza, J. V. A., Knafou, J., Oliveira, L. E. S. e., Copara, J., Gumiel, Y. B., Oliveira, L. F. A. d., Paraiso, E. C., Teodoro, D., and Barra, C. M. C. M. (2020). BioBERTpt – a Portuguese neural language model for clinical named entity recognition. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, pages 65–72. Association for Computational Linguistics.

Secretaria de Vigilância em Saúde (2024). Vigilância epidemiológica. Ministério da Saúde do Brasil. Disponível em: [link]. Acesso em: mai. 2026.

Souza, F., Nogueira, R., and Lotufo, R. (2020). BERTimbau: Pretrained BERT models for Brazilian Portuguese. In Intelligent Systems (BRACIS 2020), pages 403–417. Springer.

Tkachenko, M., Malyuk, M., Holmberg, N., Liubimov, N., and Strulov, A. (2020). Label Studio: Data labeling software. [link].

Uzuner, Ö., South, B. R., Shen, S., and DuVall, S. L. (2011). 2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text. Journal of the American Medical Informatics Association, 18(5):552–556.
Publicado
19/10/2026
FREITAS, Christian; BERTON, Lilian. EpiCorpus-BR: A Named Entity Recognition Corpus for Epidemiological Documents in Portuguese. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 150-162. DOI: https://doi.org/10.5753/stil.2026.26614.