Mitigating LLM Hallucinations in Code Generation for BDI Agents

  • Natália Mendes Goes UTFPR
  • Rafael C. Cardoso University of Aberdeen
  • Gleifer V. Alves UTFPR
  • André P. Borges UTFPR

Resumo


Este artigo apresenta o MASPY-LLM, um pipeline autônomo para geração de código projetado para mitigar as alucinações de Large Language Models (LLMs) na programação de agentes Belief-Desire-Intention (BDI) no framework Multi-Agent System for Python (MASPY). A arquitetura híbrida integra injeção de contexto, Retrieval-Augmented Generation (RAG), orquestração ator-crítico e higienização algorítmica para corrigir quebras topológicas e previnir amnésia de contexto. Para avaliar a abordagem da ferramenta, instanciou-se um estudo de caso focado no protocolo de negociação contract-net. Os resultados demonstram a capacidade do pipeline em compilar códigos sintaticamente válido e aderente às restrições do MASPY, reduzindo a curva de aprendizado na modelagem de sistemas no framework.

Referências

Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623.

Bratman, M. E. (1987). Intention, Plans, and Practical Reason. Harvard University Press, Cambridge, MA.

Cardoso, R. C. and Ferrando, A. (2021). A review of agent-based programming for multi-agent systems. Computers, 10(2):16.

Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z. (2024a). M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. In Findings of the Association for Computational Linguistics: ACL 2024, pages 2318–2335, Bangkok, Thailand. Association for Computational Linguistics.

Chen, X., Lin, M., Schärli, N., and Zhou, D. (2024b). Teaching large language models to self-debug. In International Conference on Learning Representations, volume 2024, pages 8746–8825.

Ciatto, G., Aguzzi, G., Battistini, R., Baiardi, M., Burattini, S., and Ricci, A. (2025). Exploiting genai for plan generation in bdi agents. In ECAI 2025, volume 413 of Frontiers in Artificial Intelligence and Applications, pages 3495–3502. IOS Press.

Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., and Zhang, X. (2024). Large language model based multi-agents: A survey of progress and challenges. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, pages 8048–8057. International Joint Conferences on Artificial Intelligence Organization. Survey Track.

Junior, U. G., Born, M. B., Santos, A. C., Grossmann, R. B., Facklamm, J. V., de Castilhos, V. A., Alves, B. C., and de Aguiar, M. S. (2025). Sistemas multiagente e large language model: estudo de caso utilizando as ferramentas LM Studio e LangGraph. In Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações (WESAAC), pages 250–261. SBC.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in neural information processing systems, 33:9459–9474.

Ma, X., Fang, G., and Wang, X. (2023). LLM-pruner: On the structural pruning of large language models. Advances in neural information processing systems, 36:21702–21720.

Mellado, A. L. L., Borges, A. P., and Alves, G. V. (2025). MASPY: A Python-based framework for developing BDI multi-agent systems. In Advances in Practical Applications of Agents, Multi-Agent Systems, and Computational Social Science: The PAAMS Collection, pages 216–227. Springer-Verlag, Berlin, Heidelberg.

Neres, G. G., Cardoso, R. C., Borges, A. P., and Alves, G. V. (2025). Integração do SUMO e Traffic3D com o framework de agentes MASPY. In Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações (WESAAC), pages 71–78. SBC.

Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22.

Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., et al. (2024). Chatdev: Communicative agents for software development. In Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), pages 15174–15186.

Rao, A. S. and Georgeff, M. (1995). BDI Agents: From Theory to Practice. In Proc. 1st Int. Conf. Multi-Agent Systems (ICMAS), pages 312–319, San Francisco, USA.

Shanahan, M. (2024). Talking about large language models. Communications of the ACM, 67(2):68–79.

Silva, E., Santos, F. A., Thompson, P., and dos Reis, J. C. (2025). LLM-powered conversational multi-agent cognitive system for collaborative task solving. In Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações (WESAAC), pages 59–70. SBC.

Smith, R. G. (1988). The contract net protocol: High-level communication and control in a distributed problem solver. In Readings in distributed artificial intelligence, pages 357–366. Elsevier.

Wooldridge, M. (2009). An introduction to multiagent systems. John wiley & sons.

Wu, L., Zheng, Z., Qiu, Z., Wang, H., Gu, H., Shen, T., Qin, C., Zhu, C., Zhu, H., Liu, Q., et al. (2024). A survey on large language models for recommendation. World Wide Web, 27(5):60.
Publicado
19/10/2026
GOES, Natália Mendes; CARDOSO, Rafael C.; ALVES, Gleifer V.; BORGES, André P.. Mitigating LLM Hallucinations in Code Generation for BDI Agents. In: WORKSHOP-ESCOLA DE SISTEMAS DE AGENTES, SEUS AMBIENTES E APLICAÇÕES (WESAAC), 20. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 228-239. ISSN 2326-5434. DOI: https://doi.org/10.5753/wesaac.2026.31818.