Aprendizado por Reforço para Alocação de Rins: Uma Abordagem Baseada em Policy Gradient com Dados Reais do Sistema Brasileiro de Transplantes
Resumo
Este trabalho apresenta um agente de alocação de rins baseado em Aprendizado por Reforço Profundo (DRL), visando otimizar a sobrevida no sistema brasileiro de transplantes. O modelo foi treinado e avaliado em cenários pré e durante a COVID-19, demonstrando convergência estável. Os resultados indicam que a parametrização da recompensa permite ajustar a política conforme as prioridades clínicas: a configuração C1 reduziu a taxa de mortalidade em lista de 11,32% para 8,07% (sem pandemia) e de 10,04% para 6,67% (durante a pandemia), superando o sistema atual (BRAS) em todas as métricas e diminuindo o tempo médio de espera pandémico de 476,7 para 188,9 dias. Já a configuração C2 priorizou estritamente a mortalidade em lista, reduzindo-a para os mínimos de 7,59% (sem pandemia) e 6,36% (durante a pandemia), embora com um impacto severo de aumento no tempo de espera sob condições pandémicas.
Referências
Deshpande, R. (2024). Smart match: revolutionizing organ allocation through artificial intelligence. Frontiers in Artificial Intelligence, Volume 7 - 2024.
Ding, S., Zha, D., Zhang, K., Chen, L., Jiang, X., Hu, X., and Zou, N. (2025). Fairalloc: Learning fair organ allocation policy for liver transplant. Journal of Healthcare Informatics Research.
Elalouf, A. and Pliskin, J. S. (2022). Balancing equity and efficiency in kidney allocation: An overview. Cambridge Quarterly of Healthcare Ethics, 31(3):321–332.
Jalilvand, N., Bairamzadeh, S., Tavakkoli-Moghaddam, R., and Azaron, A. (2023). A bi-objective organ transplant supply chain network with recipient priority considering carbon emission under uncertainty. Computers & Industrial Engineering, 181:109295.
Kupiec-Weglinski, J. W. (2022). Grand challenges in organ transplantation. Frontiers in Transplantation, 1:897679.
Li, H., Zhang, W., and Chen, L. (2023). Predicting kidney transplant outcomes using machine learning: A systematic review. Frontiers in Medicine, 10:1158704.
Lima, B. A., Reis, F., Alves, H., and Henriques, T. S. (2023). Equity matrix for kidney transplant allocation. Transplant Immunology, 81:101917.
Matas, A. J., Smith, D. L., Skeans, J. A., and Stewart, J. P. (2023). Optn/srtr 2022 annual data report: Kidney. American Journal of Transplantation, 23(S1):1–49.
Naqvi, S. A. A., Tennankore, K., Worthen, G., Vinson, A., and Abidi, S. S. R. (2025). A reinforcement learning framework for optimizing kidney allocation for transplant based on survival and ethical criteria. In Artificial Intelligence in Medicine, pages 323–332. Springer Nature.
Okano, C. d. S., Menezes, C. C. S. d., Brandão, F. A., Carletto, V. R., and de Souza, M. A. (2023). Analysis of the national transplant scenario in brazil. Research, Society and Development, 12(9).
Sá, G. C. B. e. and Madeira, C. A. G. (2025). Deep reinforcement learning in real-time strategy games: a systematic literature review. Applied Intelligence, 55(3):243.
Salomão Pontes, D. F., Fernandes Ferreira, G., Segev, D., Massie, A. B., Levan, M., Barbosa, A. M. P., da Rocha, N. C., and Modelli de Andrade, L. G. (2024). Regional disparities in kidney transplant allocation in brazil: A retrospective cohort study. Clinical Transplantation, 38(9):e15446.
Sutton, R. S. and Barto, A. G. (2018). Reinforcement Learning: An Introduction. MIT Press, 2nd edition.
Tang, C., Abbatematteo, B., Hu, J., Chandra, R., Martín-Martín, R., and Stone, P. (2025). Deep reinforcement learning for robotics: A survey of real-world successes. Annual Review of Control, Robotics, and Autonomous Systems, 8:153–188.
Tonelli, M., Wiebe, N., Knoll, G., and et al. (2011). Kidney transplantation compared with dialysis in clinically relevant outcomes: a systematic review. The Lancet, 378(9800):180–189.
Zhao, D., Huanshi, X., and Zhang, X. (2024). Active exploration deep reinforcement learning for continuous action space with forward prediction. International Journal of Computational Intelligence Systems, 17.
Zhu, L., Gao, W., Zhang, S., et al. (2024). Equality of healthcare resource allocation between impoverished counties and non-impoverished counties in northwest china: a longitudinal study. BMC Health Services Research, 24.
Zhu, Y. et al. (2025). Deep reinforcement learning of mobile robot navigation: A comparative analysis of value-based, policy-based and hybrid methods. Sensors, 25(11).
