Extração Clínica em Texto Livre: Uma Análise Orientada pela Complexidade Semântica
Resumo
Este trabalho investiga a extração automatizada de informações clínicas a partir de notas médicas não estruturadas em português, considerando quatro fases de complexidade semântica crescente: scores e tempos clínicos, medicamentos na alta, comorbidades e complicações. Comparamos métodos determinísticos, modelos generativos com prompt engineering, versões com fine-tuning e com reasoning. Os resultados mostram que a adequação do método depende da natureza da variável: tarefas mais simples podem ser resolvidas por abordagens léxicas, enquanto as mais ambíguas exigem modelos mais robustos. Esses achados reforçam a importância de analisar, conjuntamente, desempenho e custo computacional na estruturação de dados clínicos.
Palavras-chave:
mineração e análise de dados, recuperação de informação, mineração de textos e processamento de linguagem natural
Referências
Agrawal, M., Hegselmann, S., Lang, H., Kim, Y., and Sontag, D. (2022). Large language models are few-shot clinical information extractors. In Proceedings of the 2022 conference on empirical methods in natural language processing, pages 1998–2022.
Anschau, F., Aredes, N. D. A., and others. (2022). Cohort study protocol of the brazilian collaborative research network on covid-19: strengthening who global data. BMJ Open, 12(11).
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized language models. arXiv preprint arXiv:2305.14314.
Fox, K. A. A., Dabbous, O. H., Goldberg, R. J., Pieper, K. S., Eagle, K. A., Van de Werf, F., Avezum, Á., Goodman, S. G., Flather, M. D., Anderson, F. A., and Granger, C. B. (2006). Prediction of risk of death and myocardial infarction in the six months after presentation with acute coronary syndrome: prospective multinational observational study (GRACE). BMJ, 332(7553):1091–1094.
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
Levenshtein, V. I. (1966). Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10(8):707–710.
Meta AI (2024). The Llama 3 herd of models. arXiv preprint arXiv:2407.21783.
Morrow, D. A., Antman, E. M., Charlesworth, A., et al. (2000). A new risk score for patients with acute coronary syndromes treated with percutaneous coronary intervention. Journal of the American College of Cardiology, 43(1):89–96.
Oliveira, L. E. S. e., Peters, A. C., Da Silva, A. M. P., Gebeluca, C. P., Gumiel, Y. B., Cintho, L. M. M., Carvalho, D. R., Al Hasan, S., and Moro, C. M. C. (2022). Semclinbr-a multi-institutional and multi-specialty semantically annotated corpus for portuguese clinical nlp tasks. Journal of Biomedical Semantics, 13(1):13.
Reading Turchioe, M., Volodarskiy, A., Pathak, J., Wright, D. N., Tcheng, J. E., and Slotwiner, D. (2022). Systematic review of current natural language processing methods and applications in cardiology. Heart, 108(12):909–916.
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al. (2023). Large language models encode clinical knowledge. Nature, 620:172–180.
Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W. (2023). Large language models in medicine. Nature medicine, 29(8):1930–1940.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., and et al. (2025). Qwen3 technical report. arXiv preprint arXiv:2505.09388.
Zanotto, B. S., Beck da Silva Etges, A. P., Dal Bosco, A., et al. (2021). Stroke outcome measurements from electronic medical records: cross-sectional study on the effectiveness of neural and nonneural classifiers. JMIR Medical Informatics, 9(11):e29120.
Anschau, F., Aredes, N. D. A., and others. (2022). Cohort study protocol of the brazilian collaborative research network on covid-19: strengthening who global data. BMJ Open, 12(11).
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized language models. arXiv preprint arXiv:2305.14314.
Fox, K. A. A., Dabbous, O. H., Goldberg, R. J., Pieper, K. S., Eagle, K. A., Van de Werf, F., Avezum, Á., Goodman, S. G., Flather, M. D., Anderson, F. A., and Granger, C. B. (2006). Prediction of risk of death and myocardial infarction in the six months after presentation with acute coronary syndrome: prospective multinational observational study (GRACE). BMJ, 332(7553):1091–1094.
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
Levenshtein, V. I. (1966). Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10(8):707–710.
Meta AI (2024). The Llama 3 herd of models. arXiv preprint arXiv:2407.21783.
Morrow, D. A., Antman, E. M., Charlesworth, A., et al. (2000). A new risk score for patients with acute coronary syndromes treated with percutaneous coronary intervention. Journal of the American College of Cardiology, 43(1):89–96.
Oliveira, L. E. S. e., Peters, A. C., Da Silva, A. M. P., Gebeluca, C. P., Gumiel, Y. B., Cintho, L. M. M., Carvalho, D. R., Al Hasan, S., and Moro, C. M. C. (2022). Semclinbr-a multi-institutional and multi-specialty semantically annotated corpus for portuguese clinical nlp tasks. Journal of Biomedical Semantics, 13(1):13.
Reading Turchioe, M., Volodarskiy, A., Pathak, J., Wright, D. N., Tcheng, J. E., and Slotwiner, D. (2022). Systematic review of current natural language processing methods and applications in cardiology. Heart, 108(12):909–916.
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al. (2023). Large language models encode clinical knowledge. Nature, 620:172–180.
Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W. (2023). Large language models in medicine. Nature medicine, 29(8):1930–1940.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., and et al. (2025). Qwen3 technical report. arXiv preprint arXiv:2505.09388.
Zanotto, B. S., Beck da Silva Etges, A. P., Dal Bosco, A., et al. (2021). Stroke outcome measurements from electronic medical records: cross-sectional study on the effectiveness of neural and nonneural classifiers. JMIR Medical Informatics, 9(11):e29120.
Publicado
08/09/2026
Como Citar
PEREIRA, Antônio; PEREIRA, Lucas; MARCOLINO, Miriam Allein Zago; CARDOSO, Ricardo Bertoglio; LARA, Luciana; ETGES, Ana Paula Beck da Silva; POLANCZYK, Carisi A.; ROCHA, Leonardo.
Extração Clínica em Texto Livre: Uma Análise Orientada pela Complexidade Semântica. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 812-818.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249380.
