Multi-Agent Systems and Prompt Engineering in automating and optimizing prompt generation and refinement processes: a Systematic Literature Review

  • Allan M. Gonçalves FURG
  • Cleo Z. Billa FURG
  • Diana F. Adamatti FURG

Resumo


Esta Revisão Sistemática da Literatura investigou como Sistemas Multiagentes têm sido aplicados na Engenharia de Prompts para otimizar e automatizar a interação com Grandes Modelos de Linguagem (LLMs). O estudo foi conduzido seguindo as diretrizes PRISMA2020 e um protocolo baseado na estratégia PICO, com buscas em bases científicas consolidadas consolidadas. Inicialmente, foram encontrados 39 estudos, dos quais apenas 6 atenderam plenamente aos critérios de inclusão. Conclui-se que a integração entre sistemas multiagentes, refinamento iterativo e mecanismos estruturados de avaliação representam uma evolução para a Engenharia de prompts, contribuindo para a obtenção de respostas mais precisas e alinhadas às intenções dos usuários.

Referências

Dorsch, R., Henselmann, D., and Harth, A. (2025). Compass: A process mining-based methodology for prompt optimization of large language model agents. CEUR Workshop Proceedings, 3996:29 – 37.

Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., and Zhang, X. (2024). Large language model based multi-agents: A survey of progress and challenges. Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence Survey Track, pages 8048–8057.

Hong, M., Lee, E., Park, S., and Kim, J. (2026). Peem: Prompt engineering evaluation metrics for interpretable joint evaluation of prompts and responses in llms. IEEE Access, 14:53581–53600.

Jalori, G., Verma, P., and Arık, S. (2025). Flairr-ts – forecasting llm-agents with iterative refinement and retrieval for time series. EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2025, page 15427 – 15437.

Jason Wei, Xuezhi Wang, D. S. M. B. B. I. F. X. E. H. C. Q. V. L. and Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS ’22), page 24824–24837.

Kitchenham, B. A. and Charters, S. (2007). Guidelines for performing systematic literature reviews in software engineering – version 2.3. Technical report, Keele/Staffs-UK and Durham-UK.

Nakagawa, E. Y., Scannavino, K. R. F., Fabbri, S. C. P. F., and Ferrari, F. C. (2017). Revisão Sistemática da Literatura em Engenharia de Software: teoria e prática. Elsevier.

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., McGuinness, L. A., Stewart, L. A., Thomas, J., Tricco, A. C., Welch, V. A., Whiting, P., and Moher, D. (2021). The prisma 2020 statement: an updated guideline for reporting systematic reviews. BMJ, 372.

Pai, M., McCulloch, M., Gorman, J. D., Pai, N., Enanoria, W., Kennedy, G., Tharyan, P., and Colford Jr., J. M. (2004). Systematic reviews and meta-analyses: An illustrated, step-by-step guide. National Medical Journal of India, 17(2):86–95.

Purpura, A., Wang, L., Badyal, S., Beaufrand, E., and Faulkner, A. (2026). Enhancing llm instruction following: An evaluation-driven multi-agentic workflow for prompt instructions optimization.

Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., and Chadha, A. (2025). A systematic survey of prompt engineering in large language models: Techniques and applications.

Spiess, C., Vaziri, M., Mandel, L., and Hirzel, M. (2025). Autopdl: Automatic prompt optimization for llm agents. Proceedings of Machine Learning Research, 293.

Tian, J., Fard, P., Cagan, C., et al. (2026). An autonomous agentic workflow for clinical detection of cognitive concerns using large language models. npj Digital Medicine, 9(1):51.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2023). Attention is all you need.

Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. (2026). A survey of large language models. Front. Comput. Sci. 20.
Publicado
19/10/2026
GONÇALVES, Allan M.; BILLA, Cleo Z.; ADAMATTI, Diana F.. Multi-Agent Systems and Prompt Engineering in automating and optimizing prompt generation and refinement processes: a Systematic Literature Review. In: WORKSHOP-ESCOLA DE SISTEMAS DE AGENTES, SEUS AMBIENTES E APLICAÇÕES (WESAAC), 20. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 134-145. ISSN 2326-5434. DOI: https://doi.org/10.5753/wesaac.2026.31380.