Stagent: A Finite-State Machine Architecture for Traceable LLM Agents

  • Lucas T. Bandeira UNICAMP
  • Daniel Alves da Rocha UNICAMP
  • Eryck Silva UNICAMP
  • Julio C. dos Reis UNICAMP

Resumo


Autonomous agents extend LLMs to multi-step tasks by combining reasoning, intermediate decisions, and tool use. However, many LLM-based workflows remain implicit execution chains, making failure localization, recovery control, and intermediate evaluation difficult. This investigation presents Stagent, a finite-state machine architecture that models agent execution as explicit states and transitions, enabling workflow steps to be inspected and redirected after failures. For experimental purposes, we specialize Stagent for tabular document diagnosis as CsvStagent, obtaining valid Finite State Machine (FSM) traces in all executions and reaching 0.963 task success. Our findings show that deterministic diagnostic context and FSM-based verification support traceable diagnosis and future self-healing.

Referências

Bahroun, Z., Anane, C., Ahmed, V., and Zacca, A. (2023). Transforming education: A comprehensive review of generative artificial intelligence in educational settings through bibliometric and content analysis. Sustainability, 15(17):12983.

Bendinelli, T., Dox, A., and Holz, C. (2025). Exploring LLM agents for cleaning tabular machine learning datasets. In ICLR 2025 Workshop on Foundation Models in the Wild.

Cho, N., Fielding, K., Watson, W., Ganesh, S., and Veloso, M. (2026). TASER: Table agents for schema-guided extraction and recommendation. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), pages 226–252. Association for Computational Linguistics.

Choudhury, M. R., Iyengar, A. I. K. N., Siingh, S., Puranam, S., and Gupta, V. (2025). TABARD: A novel benchmark for tabular anomaly analysis, reasoning and detection. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V., editors, Findings of the Association for Computational Linguistics: EMNLP 2025, pages 21783–21817, Suzhou, China. Association for Computational Linguistics.

Gomaa, H. (2011). Finite State Machines, page 151–176. Cambridge University Press.

Jo, E., Epstein, D. A., Jung, H., and Kim, Y.-H. (2023). Understanding the benefits and challenges of deploying conversational AI leveraging large language models for public health intervention. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–16. ACM.

Konolige, K. and Nilsson, N. J. (1980). Multiple-agent planning systems. In AAAI, volume 80, pages 138–142.

Lampert, L., Hübner, J., and Zatelli, M. (2019). Tolerância a faltas em sistemas multiagentes multidimensionais. In Anais do XIII Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações, pages 230–235, Porto Alegre, RS, Brasil. SBC.

Liu, J., Shuai, J., and Li, X. (2024). State machine of thoughts: Leveraging past reasoning trajectories for enhancing problem solving. arXiv:2312.17445.

Mohamed, A. H., Santos, F. A., and dos Reis, J. C. (2025). Enhancing llm agent effectiveness via reflective multi-agent system. In Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações (WESAAC), pages 206–217. SBC.

Silva, E., Santos, F. A., Thompson, P., and dos Reis, J. C. (2025). Llm-powered conversational multi-agent cognitive system for collaborative task solving. In Workshop-Escola de Sistemas de Agentes, seus Ambientes e Aplicações (WESAAC), pages 59–70. SBC.

Sumers, T., Yao, S., Narasimhan, K., and Griffiths, T. (2024). Cognitive architectures for language agents. Transactions on Machine Learning Research.

Vitagliano, G., Hameed, M., Jiang, L., Reisener, L., Wu, E., and Naumann, F. (2023). Pollock: A data loading benchmark. Proc. VLDB Endow., 16(8):1870–1882.

Wu, Y., Yue, T., Zhang, S., Wang, C., and Wu, Q. (2024). Stateflow: Enhancing llm task-solving through state-driven workflows. arXiv:2403.11322.

Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2023a). Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems, 36:11809–11822.

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2023b). React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR).

Zhou, J., Chen, J., Lu, Q., Zhao, D., and Zhu, L. (2025). Shielda: Structured handling of exceptions in llm-driven agentic workflows. arXiv:2508.07935.
Publicado
19/10/2026
BANDEIRA, Lucas T.; ROCHA, Daniel Alves da; SILVA, Eryck; REIS, Julio C. dos. Stagent: A Finite-State Machine Architecture for Traceable LLM Agents. In: WORKSHOP-ESCOLA DE SISTEMAS DE AGENTES, SEUS AMBIENTES E APLICAÇÕES (WESAAC), 20. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 276-287. ISSN 2326-5434. DOI: https://doi.org/10.5753/wesaac.2026.32015.