Podemos Confiar em LLMs na Recomendação Clínica? Uma Avaliação Preliminar sobre Viés e Confiabilidade
Resumo
Os Modelos de Linguagem de Larga Escala (LLMs) vêm sendo aplicados na área da saúde, mas suscitam preocupações quanto a vieses e à confiabilidade clínica. Este artigo propõe um workflow para a identificação de vieses em recomendações clínicas com base na avaliação contrafactual, mantendo o quadro clínico fixo e variando atributos como sexo e raça. Os experimentos consideraram 72 inferências por modelo, combinando seis cenários-base, seis perfis contrafactuais e dois tipos de prompt. As respostas foram avaliadas quanto à aderência às diretrizes clínicas e classificadas como corretas, com viés potencial ou como erro clínico. Os resultados indicam que cerca de 50% das recomendações apresentam problemas, com maior incidência associada à raça.
Palavras-chave:
Modelos De Linguagem, Viés Algoritmico, Avaliacao Contrafactual, Recomendação Clínica, Equidade Em Saude, Farmacologia
Referências
Amorim, A., Assis, G., et al. (2026). When prompts know the story: Improving llm detection of online hate speech. Journal of Information and Data Management.
Brandão, A. A. et al. (2025). Diretriz brasileira de hipertensão arterial. Arq. Bras. de Cardiologia. Soc. Bras. de Cardiologia, Soc. Bras. de Hipertensão e Soc. Bras. de Nefrologia.
da Silva, D. C. A. et al. (2025). Analysis of the effectiveness of llms in handwritten essay recognition and assessment. In Proc. of the 17th ICAART, pages 776–785.
Gilson, A. et al. (2023). How well does chatgpt perform on the usmle? JMIR Medical Edu.
Han, Y. and Tao, J. (2024). Revolutionizing pharma: Unveiling the AI and LLM trends in the pharmaceutical industry. CoRR, abs/2401.10273.
Holdcroft, A. (2007). Gender bias in research: how does it affect evidence based medicine? Journal of the Royal Society of Medicine, 100(1):2–3.
Jones, D. W. et al. (2025). 2025 guideline for the prevention, detection, evaluation, and management of high blood pressure in adults. Hypertension, 82:e212–e316.
Kusner, M. J., Loftus, J., Russell, C., and Silva, R. (2017). Counterfactual fairness. In Advances in Neural Information Processing Systems.
Li, L. et al. (2026). Llm use for mental health: Crowdsourcing users’ sentiment-based perspectives and values from social discussions. In ACM WWW’26, page 9687–9698.
Min. Saúde/BR (2024). Protocolo clínico e diretrizes terapêuticas da dor crônica: documento resumido. Portaria Conjunta SAES/SAPS/SECTICS/MS nº 01, de 22 de agosto de 2024.
Min. Saúde/BR (2025). Protocolo clínico e diretrizes terapêuticas da hipertensão arterial sistêmica. Portaria SECTICS/MS nº 49, de 23 de julho de 2025.
Minson, F. P. et al. (2011). II Consenso Nacional de Dor Oncológica, volume 1. grupo editorial moreira jr.
NICE (2020). Low back pain and sciatica in over 16s: Assessment and management. NICE guideline NG59.
Obermeyer, Z., Powers, B., Vogeli, C., and Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453.
Rostami, M., Hossain, K. S. M. T., and Hawamdeh, S. (2025). Detecting health misinformation by leveraging llm models and debunk list. In Proc. of the ACM/IEEE CHASE, page 341–346, New York, NY, USA. ACM.
Shaw, L. J., Bugiardini, R., and Merz, C. N. B. (2009). Women and ischemic heart disease: Evolving knowledge. Journal of the American College of Cardiology, 54(17):1561–1575.
Singhal, K. et al. (2023). Large language models encode clinical knowledge. Nature.
Zhao, W. X. et al. (2023). A survey of large language models. arXiv:2303.18223, 1(2).
Brandão, A. A. et al. (2025). Diretriz brasileira de hipertensão arterial. Arq. Bras. de Cardiologia. Soc. Bras. de Cardiologia, Soc. Bras. de Hipertensão e Soc. Bras. de Nefrologia.
da Silva, D. C. A. et al. (2025). Analysis of the effectiveness of llms in handwritten essay recognition and assessment. In Proc. of the 17th ICAART, pages 776–785.
Gilson, A. et al. (2023). How well does chatgpt perform on the usmle? JMIR Medical Edu.
Han, Y. and Tao, J. (2024). Revolutionizing pharma: Unveiling the AI and LLM trends in the pharmaceutical industry. CoRR, abs/2401.10273.
Holdcroft, A. (2007). Gender bias in research: how does it affect evidence based medicine? Journal of the Royal Society of Medicine, 100(1):2–3.
Jones, D. W. et al. (2025). 2025 guideline for the prevention, detection, evaluation, and management of high blood pressure in adults. Hypertension, 82:e212–e316.
Kusner, M. J., Loftus, J., Russell, C., and Silva, R. (2017). Counterfactual fairness. In Advances in Neural Information Processing Systems.
Li, L. et al. (2026). Llm use for mental health: Crowdsourcing users’ sentiment-based perspectives and values from social discussions. In ACM WWW’26, page 9687–9698.
Min. Saúde/BR (2024). Protocolo clínico e diretrizes terapêuticas da dor crônica: documento resumido. Portaria Conjunta SAES/SAPS/SECTICS/MS nº 01, de 22 de agosto de 2024.
Min. Saúde/BR (2025). Protocolo clínico e diretrizes terapêuticas da hipertensão arterial sistêmica. Portaria SECTICS/MS nº 49, de 23 de julho de 2025.
Minson, F. P. et al. (2011). II Consenso Nacional de Dor Oncológica, volume 1. grupo editorial moreira jr.
NICE (2020). Low back pain and sciatica in over 16s: Assessment and management. NICE guideline NG59.
Obermeyer, Z., Powers, B., Vogeli, C., and Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453.
Rostami, M., Hossain, K. S. M. T., and Hawamdeh, S. (2025). Detecting health misinformation by leveraging llm models and debunk list. In Proc. of the ACM/IEEE CHASE, page 341–346, New York, NY, USA. ACM.
Shaw, L. J., Bugiardini, R., and Merz, C. N. B. (2009). Women and ischemic heart disease: Evolving knowledge. Journal of the American College of Cardiology, 54(17):1561–1575.
Singhal, K. et al. (2023). Large language models encode clinical knowledge. Nature.
Zhao, W. X. et al. (2023). A survey of large language models. arXiv:2303.18223, 1(2).
Publicado
08/09/2026
Como Citar
FERRARI, Camila; DE OLIVEIRA, Daniel.
Podemos Confiar em LLMs na Recomendação Clínica? Uma Avaliação Preliminar sobre Viés e Confiabilidade. In: BRAZILIAN E-SCIENCE WORKSHOP (BRESCI), 20. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 57-64.
ISSN 2763-8774.
DOI: https://doi.org/10.5753/bresci.2026.249437.
