Leakage-Aware Evaluation of Brazilian Clinical Notes for ICU Mortality Prediction: A Study on the BRATECA Dataset
Resumo
Early prediction of ICU mortality can support clinical decisions, but retrospective notes can contain future-event cues that cause leakage in clinical NLP. We evaluate leakage-aware modeling of Brazilian Portuguese ICU notes for mortality prediction using BRATECA. We compare 24-hour restriction, regex and LLM leakage auditing, neural and TF-IDF representations, and unrestricted-note baselines. In the 24-hour setting, models achieved AUROC 0.78-0.83; unrestricted notes reached around 0.95, indicating outcome-related information in late documentation. Regex found leakage signals in 21.13% of the notes, and the LLM audit identified additional semantic cases. Results highlight the need for temporal control and leakage auditing in clinical NLP.Referências
Alsentzer, E., Murphy, J., Boag, W., et al. (2019). Publicly available clinical bert embeddings. In Proceedings of the Clinical NLP Workshop.
Boll, H. O., Amirahmadi, A., Ghazani, M. M., de Morais, W. O., de Freitas, E. P., Soliman, A., Etminani, F., Byttner, S., and Recamonde-Mendoza, M. (2024). Graph neural networks for clinical risk prediction based on electronic health records: A survey. Journal of Biomedical Informatics.
Chen, Z., Hernández Cano, A., Romanou, A., Bonnet, A., Matoba, K., Salvi, F., Pagliardini, M., et al. (2023). Meditron-70b: Scaling medical pretraining for large language models.
Chiavegatto Filho, A. D. P., Batista, A. F., and Santos, H. G. (2021). Data leakage in health outcomes prediction with machine learning. comment on ”prediction of incident hypertension within the next year: Prospective study using statewide electronic health records and machine learning”. Journal of Medical Internet Research, 23(2).
Choi, M. H., Kim, D., Choi, E. J., Jung, Y. J., Choi, Y. J., Cho, J. H., and Jeong, S. H. (2022). Mortality prediction of patients in intensive care units using machine learning algorithms based on electronic health records. Scientific Reports, 12(1):7180.
Consoli, B. S., dos Santos, H. D. P., Ulbrich, A. H. D. P. S., Vieira, R., and Bordini, R. H. (2022). Brateca (brazilian tertiary care dataset): a clinical information dataset for the portuguese language. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (LREC 2022), Marseille, France. European Language Resources Association (ELRA).
Costa-jussà, M. R., Cross, J., Červeňansky, F., et al. (2022). No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672.
Davis, S. E., Matheny, M. E., Balu, S., and Sendak, M. P. (2023). A framework for understanding label leakage in machine learning for health care. Journal of the American Medical Informatics Association: JAMIA, 31(1):274.
Dias, H. and Ulbrich, A. H. D. P. d. (2022). BRATECA (Brazilian Tertiary Care Dataset): a Clinical Information Dataset for the Portuguese Language. PhysioNet. Version 1.1.
Garriga, R., Buda, T. S., Guerreiro, J., Omaña Iglesias, J., Estella Aguerri, I., and Matić, A. (2023). Combining clinical notes with structured electronic health records enhances the prediction of mental health crises. Cell Reports Medicine, 4(11):101260.
Goldberger, A. L., Amaral, L. A., Glass, L., Hausdorff, J. M., Ivanov, P. C., Mark, R. G., Mietus, J. E., Moody, G. B., Peng, C.-K., and Stanley, H. E. (2000). Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals. circulation, 101(23):e215–e220.
Johnson, A. E., Pollard, T. J., Shen, L., et al. (2016). Mimic-iii, a freely accessible critical care database. Scientific Data, 3:160035.
Knaus, W. A., Draper, E. A., Wagner, D. P., and Zimmerman, J. E. (1985). Apache ii: a severity of disease classification system. Critical Care Medicine, 13(10):818–829.
Le Gall, J.-R., Lemeshow, S., and Saulnier, F. (1993). A new simplified acute physiology score (saps ii) based on a european/north american multicenter study. JAMA, 270(24):2957–2963.
Lee, J., Yoon, W., Kim, S., et al. (2020). Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
Nogueira, R., Pires, R., Abonizio, H., et al. (2023). Sabiá: Portuguese large language models.
Olang, O., Mohseni, S., Shahabinezhad, A., Hamidianshirazi, Y., Goli, A., Abolghasemian, M., Shafiee, M. A., Aarabi, M., Alavinia, M., and Shaker, P. (2025). Artificial intelligence-based models for prediction of mortality in ICU patients: a scoping review. Journal of Intensive Care Medicine, 40(12):1240–1246.
OpenAI (2023). Gpt-4 technical report.
Seinen, T. M., Kors, J. A., van Mulligen, E. M., and Rijnbeek, P. R. (2025). Using structured codes and free-text notes to measure information complementarity in electronic health records: Feasibility and validation study. Journal of Medical Internet Research, 27:e66910.
Tavabi, N., Singh, M., Pruneski, J., and Kiapour, A. M. (2024). Systematic evaluation of common natural language processing techniques to codify clinical notes. Plos one, 19(3):e0298892.
Vagliano, I., Dormosh, N., Rios, M., Luik, T. T., Buonocore, T., Elbers, P. W., Dongelmans, D. A., Schut, M. C., and Abu-Hanna, A. (2023). Prognostic models of in-hospital mortality of intensive care patients using neural representation of unstructured text: A systematic review and critical appraisal. Journal of Biomedical Informatics, 146:104504.
van Aken, B., Papaioannou, J.-M., Mayrdorfer, M., Budde, K., Gers, F., and Loeser, A. (2021). Clinical outcome prediction from admission notes using self-supervised knowledge integration. In Merlo, P., Tiedemann, J., and Tsarfaty, R., editors, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 881–893, Online. Association for Computational Linguistics.
Vincent, J.-L., Moreno, R., Takala, J., Willatts, S., De Mendonça, A., Bruining, H., Reinhart, K., Suter, P. M., and Thijs, L. G. (1996). The sofa (sepsis-related organ failure assessment) score to describe organ dysfunction/failure. Intensive Care Medicine, 22:707–710.
Boll, H. O., Amirahmadi, A., Ghazani, M. M., de Morais, W. O., de Freitas, E. P., Soliman, A., Etminani, F., Byttner, S., and Recamonde-Mendoza, M. (2024). Graph neural networks for clinical risk prediction based on electronic health records: A survey. Journal of Biomedical Informatics.
Chen, Z., Hernández Cano, A., Romanou, A., Bonnet, A., Matoba, K., Salvi, F., Pagliardini, M., et al. (2023). Meditron-70b: Scaling medical pretraining for large language models.
Chiavegatto Filho, A. D. P., Batista, A. F., and Santos, H. G. (2021). Data leakage in health outcomes prediction with machine learning. comment on ”prediction of incident hypertension within the next year: Prospective study using statewide electronic health records and machine learning”. Journal of Medical Internet Research, 23(2).
Choi, M. H., Kim, D., Choi, E. J., Jung, Y. J., Choi, Y. J., Cho, J. H., and Jeong, S. H. (2022). Mortality prediction of patients in intensive care units using machine learning algorithms based on electronic health records. Scientific Reports, 12(1):7180.
Consoli, B. S., dos Santos, H. D. P., Ulbrich, A. H. D. P. S., Vieira, R., and Bordini, R. H. (2022). Brateca (brazilian tertiary care dataset): a clinical information dataset for the portuguese language. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (LREC 2022), Marseille, France. European Language Resources Association (ELRA).
Costa-jussà, M. R., Cross, J., Červeňansky, F., et al. (2022). No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672.
Davis, S. E., Matheny, M. E., Balu, S., and Sendak, M. P. (2023). A framework for understanding label leakage in machine learning for health care. Journal of the American Medical Informatics Association: JAMIA, 31(1):274.
Dias, H. and Ulbrich, A. H. D. P. d. (2022). BRATECA (Brazilian Tertiary Care Dataset): a Clinical Information Dataset for the Portuguese Language. PhysioNet. Version 1.1.
Garriga, R., Buda, T. S., Guerreiro, J., Omaña Iglesias, J., Estella Aguerri, I., and Matić, A. (2023). Combining clinical notes with structured electronic health records enhances the prediction of mental health crises. Cell Reports Medicine, 4(11):101260.
Goldberger, A. L., Amaral, L. A., Glass, L., Hausdorff, J. M., Ivanov, P. C., Mark, R. G., Mietus, J. E., Moody, G. B., Peng, C.-K., and Stanley, H. E. (2000). Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals. circulation, 101(23):e215–e220.
Johnson, A. E., Pollard, T. J., Shen, L., et al. (2016). Mimic-iii, a freely accessible critical care database. Scientific Data, 3:160035.
Knaus, W. A., Draper, E. A., Wagner, D. P., and Zimmerman, J. E. (1985). Apache ii: a severity of disease classification system. Critical Care Medicine, 13(10):818–829.
Le Gall, J.-R., Lemeshow, S., and Saulnier, F. (1993). A new simplified acute physiology score (saps ii) based on a european/north american multicenter study. JAMA, 270(24):2957–2963.
Lee, J., Yoon, W., Kim, S., et al. (2020). Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
Nogueira, R., Pires, R., Abonizio, H., et al. (2023). Sabiá: Portuguese large language models.
Olang, O., Mohseni, S., Shahabinezhad, A., Hamidianshirazi, Y., Goli, A., Abolghasemian, M., Shafiee, M. A., Aarabi, M., Alavinia, M., and Shaker, P. (2025). Artificial intelligence-based models for prediction of mortality in ICU patients: a scoping review. Journal of Intensive Care Medicine, 40(12):1240–1246.
OpenAI (2023). Gpt-4 technical report.
Seinen, T. M., Kors, J. A., van Mulligen, E. M., and Rijnbeek, P. R. (2025). Using structured codes and free-text notes to measure information complementarity in electronic health records: Feasibility and validation study. Journal of Medical Internet Research, 27:e66910.
Tavabi, N., Singh, M., Pruneski, J., and Kiapour, A. M. (2024). Systematic evaluation of common natural language processing techniques to codify clinical notes. Plos one, 19(3):e0298892.
Vagliano, I., Dormosh, N., Rios, M., Luik, T. T., Buonocore, T., Elbers, P. W., Dongelmans, D. A., Schut, M. C., and Abu-Hanna, A. (2023). Prognostic models of in-hospital mortality of intensive care patients using neural representation of unstructured text: A systematic review and critical appraisal. Journal of Biomedical Informatics, 146:104504.
van Aken, B., Papaioannou, J.-M., Mayrdorfer, M., Budde, K., Gers, F., and Loeser, A. (2021). Clinical outcome prediction from admission notes using self-supervised knowledge integration. In Merlo, P., Tiedemann, J., and Tsarfaty, R., editors, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 881–893, Online. Association for Computational Linguistics.
Vincent, J.-L., Moreno, R., Takala, J., Willatts, S., De Mendonça, A., Bruining, H., Reinhart, K., Suter, P. M., and Thijs, L. G. (1996). The sofa (sepsis-related organ failure assessment) score to describe organ dysfunction/failure. Intensive Care Medicine, 22:707–710.
Publicado
19/10/2026
Como Citar
ALVES, Felipe André Bach; BOLL, Heloísa Oss; RECAMONDE-MENDOZA, Mariana.
Leakage-Aware Evaluation of Brazilian Clinical Notes for ICU Mortality Prediction: A Study on the BRATECA Dataset. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 14-25.
DOI: https://doi.org/10.5753/stil.2026.26535.
