An Automated Approach for Verifying Learned Policies of Q-Learning Agents in NetLogo

  • Mateus Rissardi UDESC
  • Fernando Santos UDESC

Resumo


O uso de algoritmos de Aprendizado por Reforço (RL) em Modelos Baseados em Agentes (MBA) é de grande interesse prático para simular ambientes complexos. Uma extensão da ferramenta NetLogo está disponível para simplificar o uso de RL em MBA. Apesar de seus benefícios, a verificação do aprendizado dos agentes nessa extensão permanece um desafio, pois exige inspeções manuais nas estruturas de dados utilizadas pelos algoritmos de RL. Este artigo propõe um verificador automatizado da política aprendida pelo agente. O verificador compara a política aprendida pelo agente com uma política desejada, informada pelo desenvolvedor, e indica se há divergência. O uso do verificador é demonstrado por meio dos MBAs Cliff Walking e Labirinto.

Referências

Bazzanella, E., Barros, M., and Santos, F. (2024). Refatoração da extensão netlogo de aprendizagem por reforço para integração com a biblioteca burlap. In Anais do XVIII Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações, pages 51–62, Porto Alegre, RS, Brasil. SBC.

Bazzanella, E. and Santos, F. (2021). Does a q-learning netlogo extension simplify the development of agent-based simulations? In Anais do XV Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações, pages 1–12, Porto Alegre, RS, Brasil. SBC.

Bernardo, P. C. and Kon, F. (2011). Padrões de testes automatizados. Tese de doutorado, Universidade de São Paulo (USP).

Chen, M., Beutel, A., Covington, P., Jain, S., Belletti, F., and Chi, E. H. (2019). Top-k off-policy correction for a reinforce recommender system. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining, pages 456–464, New York, NY, USA. Association for Computing Machinery.

Kaelbling, L. P., Littman, M. L., and Moore, A. W. (1996). Reinforcement learning: A survey. Journal of artificial intelligence research, 4:237–285.

Kons, K. (2019). Biblioteca Q-Learning para desenvolvimento de simulações com agentes na plataforma NetLogo. Trabalho de conclusão de curso, Universidade do Estado de Santa Catarina (UDESC).

Macal, C. and North, M. (2014). Introductory tutorial: Agent-based modeling and simulation. In Proceedings of the winter simulation conference 2014, pages 6–20. IEEE.

MacGlashan, J., Loftin, R., Littman, M. L., and Roberts, D. L. (2018). BURLAP: Brown-UMBC Reinforcement Learning and Planning. burlap.cs.brown.edu.

[link] Mazouni, Q., Spieker, H., Gotlieb, A., and Acher, M. (2024). Testing for fault diversity in reinforcement learning. In Proceedings of the 5th ACM/IEEE International Conference on Automation of Software Test (AST 2024), pages 136–146.

Nunes, R. H., de Jesus Toledo, B., Bello, C. C. C., de Andrade, G. O., Mariano, G. T., and Marques, R. G. (2023). Inteligência artificial e aprendizagem por reforço. Revista CBTecLE, 7(2):210–225.

Polydoros, A. S. and Nalpantidis, L. (2017). Survey of model-based reinforcement learning: Applications on robotics. Journal of Intelligent & Robotic Systems, 86(2):153–173.

Pugh, J. K., Soros, L. B., and Stanley, K. O. (2016). Quality diversity: A new frontier for evolutionary computation. Frontiers in Robotics and AI, 3:40.

Ris-Ala, R. (2023). Fundamentos de Aprendizagem por Reforço. Ed. do Autor, Rio de Janeiro.

Russell, S. J. and Norvig, P. (2022). Inteligência Artificial: Uma Abordagem Moderna. GEN LTC, Rio de Janeiro, 4 edition.

Sezgin, E. and Özkan, S. (2013). A systematic literature review on health recommender systems. In 2013 E-Health and Bioengineering Conference (EHB), pages 1–4. IEEE.

Silva, C. R., Mesquita, O., Araújo, E., and Mata, A. S. (2025). Uso da modelagem baseada em agentes no estudo de sistemas complexos. Revista Brasileira de Ensino de Física, 47:e20240464.

Sommerville, I. (2018). Engenharia de Software. Pearson, 10 edition.

Sunba, A., Hassine, J., and Ahmed, M. (2026). Testing reinforcement learning systems: A comprehensive review. Journal of Systems and Software, 231:112563.

Sutton, R. S., Barto, A. G., and Barto, A. (1998). Reinforcement learning: An introduction, volume 1. MIT press Cambridge.

Tappler, M., Córdoba, F. C., Aichernig, B. K., and Könighofer, B. (2022). Search-based testing of reinforcement learning. arXiv preprint arXiv:2205.04887.

Watkins, C. J. and Dayan, P. (1992). Q-learning. Machine learning, 8(3):279–292.

Wilensky, U. (1999). NetLogo. Center for Connected Learning and Computer-Based Modeling, Northwestern University. Evanston, IL. [link].
Publicado
19/10/2026
RISSARDI, Mateus; SANTOS, Fernando. An Automated Approach for Verifying Learned Policies of Q-Learning Agents in NetLogo. In: WORKSHOP-ESCOLA DE SISTEMAS DE AGENTES, SEUS AMBIENTES E APLICAÇÕES (WESAAC), 20. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 264-275. ISSN 2326-5434. DOI: https://doi.org/10.5753/wesaac.2026.31982.