Local Anonymization of Free-Text Robbery Police Reports in Brazilian Portuguese: a Safety-Recall-Oriented Multi-Model Benchmark
Resumo
Brazilian public-safety agencies release structured crime statistics but withhold the free-text narrative field of police reports, which is rich in modus operandi and contextual information, because it contains personal data protected by the Brazilian General Data Protection Law (LGPD). We address this deadlock for robbery reports with a fully local (on-premise) two-stage anonymization pipeline: a deterministic regular-expression stage for structured personally identifiable information (PII) and a pluggable semantic layer (dedicated NER, local LLMs served by Ollama, union and chained hybrids, and Microsoft Presidio as a system baseline). We contribute (i) an expert-annotated gold standard of 996 robbery narratives, released with the source code; (ii) an entity-level evaluation protocol whose headline metric is safety recall (how much PII actually did not leak), complemented by a re-identification-criticality analysis and bootstrap confidence intervals; and (iii) an empirical benchmark of nine engine configurations quantifying the safety–utility trade-off. The union of spaCy and qwen3:8b over the regex stage reaches the highest macro safety recall (0.986, 95% CI [0.972; 0.994]; micro 0.973) and is the most conservative (highest-recall) option, while qwen3:8b alone gives the best overall balance (safety recall 0.966, precision 0.80, micro-F1 0.826, text retention 96%) and is our recommendation for general use. In this configuration, direct identifiers (names, documents, phones, e-mails) are removed at 99–100%, with residual leakage concentrated in location quasi-identifiers. Safety recall measures PII removal, not residual reidentification risk: these figures bound leakage rather than certify anonymity. We recommend qwen3:8b for general supervised anonymization, and the union when leakage must be minimized, and discuss conditions for replication across Brazilian police forces.
Referências
Brasil. Lei n. 13.709, de 14 de agosto de 2018. Lei Geral de Proteção de Dados Pessoais (LGPD). [link], 2018.
Dernoncourt, F., Lee, J. Y., Uzuner, Ö., and Szolovits, P. De-identification of Patient Notes with Recurrent Neural Networks. Journal of the American Medical Informatics Association 24 (3): 596–606, 2017.
Efron, B. and Tibshirani, R. J. An Introduction to the Bootstrap. Chapman & Hall, New York, 1993.
FBSP. Anuário Brasileiro de Segurança Pública 2025. Fórum Brasileiro de Segurança Pública (FBSP), 19. ed., São Paulo, 2025.
Luz de Araujo, P. H., de Campos, T. E., de Oliveira, R. R. R., Stauffer, M., Couto, S., and Bermejo, P. LeNER-Br: A Dataset for Named Entity Recognition in Brazilian Legal Text. In Computational Processing of the Portuguese Language – PROPOR 2018. Lecture Notes in Computer Science, vol. 11122. Springer, Cham, pp. 313–323, 2018.
Oliveira, L. E. S. e. et al. SemClinBr – a Multi-Institutional and Multi-Specialty Semantically Annotated Corpus for Portuguese Clinical NLP Tasks. Journal of Biomedical Semantics 13 (1): 13, 2022.
Schiezaro, M., Rosa, G., Campos, B. A. G., and Pedrini, H. Guardians of the Data: NER and LLMs for Effective Medical Record Anonymization in Brazilian Portuguese. Frontiers in Public Health vol. 13, pp. 1717303, 2026.
Schneider, E. T. R., de Souza, J. V. A., Knafou, J., e Oliveira, L. E. S., Copara, J., Gumiel, Y. B., de Oliveira, L. F. A., Paraiso, E. C., Teodoro, D., and Barra, C. M. C. M. BioBERTpt – A Portuguese Neural Language Model for Clinical Named Entity Recognition. In Proceedings of the 3rd Clinical Natural Language Processing Workshop. Association for Computational Linguistics, Online, pp. 65–72, 2020.
Souza, F., Nogueira, R., and Lotufo, R. BERTimbau: Pretrained BERT Models for Brazilian Portuguese. In Intelligent Systems – BRACIS 2020. Lecture Notes in Computer Science, vol. 12319. Springer, Cham, pp. 403–417, 2020.
SSP-SP. Estatísticas de Criminalidade. [link], 2024.
Sweeney, L. k-Anonymity: A Model for Protecting Privacy. International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems 10 (5): 557–570, 2002.
Uzuner, Ö., Luo, Y., and Szolovits, P. Evaluating the State-of-the-Art in Automatic De-identification. Journal of the American Medical Informatics Association 14 (5): 550–563, 2007.
