STAB: A Complexity-Ladder Framework for Evaluating Local LLMs as Autonomous Agents

  • João Conrado V. Nogueira

Resumo


O uso de Large Language Models (LLMs) em sistemas agênticos locais enfrenta desafios críticos de consistência e confiabilidade, especialmente em hardware de consumo. Este trabalho apresenta o protocolo STAB, uma metodologia de avaliação em escada de complexidade (S1→S4) combinada a um protocolo de estabilidade temporal para diagnosticar falhas em agentes on-premise. Avaliamos oito modelos de quatro famílias (Qwen, Gemma, Llama e Mistral) em servidor com GPU NVIDIA RTX 3060 (12 GB VRAM). Os resultados revelam que as falhas agênticas originam-se predominantemente na camada de instrução e decisão contextual, e não na capacidade bruta dos modelos. Identificamos o fenômeno de tool-skipping por nomeação semântica (H12) e demonstramos que conformidade de formato e acurácia semântica são dimensões ortogonais (H10). O modelo gemma4:e4b destacou-se como referência de maior confiabilidade, mantendo desempenho perfeito em todos os níveis de complexidade.

Palavras-chave: LLMs Locais, Agentes Autônomos, Tool Calling, Instruction Following, Small Language Models, Benchmark On-Premise

Referências

Cemri, M. et al. Why do multi-agent LLM systems fail? arXiv preprint arXiv:2503.13657, 2025.

Geng, S., Josifoski, M., Peyrard, M., and West, R. Grammar-constrained decoding for structured NLP tasks without finetuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023). Association for Computational Linguistics, Singapore, 2023.

Jhandi, P. et al. Small language models for efficient agentic tool calling: Outperforming large models with targeted fine-tuning. In Proceedings of the 40th AAAI Conference on Artificial Intelligence. AAAI Press, Singapore, 2026.

Kavathekar, I. et al. Small models, big tasks: An exploratory empirical study on small language models for function calling. In Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering (EASE 2025). ACM, Istanbul, Turkey, 2025.

Liu, X. et al. AgentBench: Evaluating LLMs as agents. In Proceedings of the 12th International Conference on Learning Representations (ICLR 2024). OpenReview.net, Vienna, Austria, 2024.

Murthy, R. et al. KCIF: Knowledge-conditioned instruction following. arXiv preprint arXiv:2410.12972, 2024.

Patil, S. G. et al. The Berkeley function calling leaderboard (BFCL): From tool use to agentic evaluation of large language models. In Proceedings of the 42nd International Conference on Machine Learning (ICML 2025). PMLR, Vancouver, Canada, 2025.

Pham, N. T. et al. SLM-Bench: A comprehensive benchmark of small language models on environmental impacts. In Findings of the Association for Computational Linguistics: EMNLP 2025. Association for Computational Linguistics, Suzhou, China, 2025.

Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023). Curran Associates, Inc., New Orleans, USA, 2023.

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. ReAct: Synergizing reasoning and acting in language models. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023). OpenReview.net, Kigali, Rwanda, 2023.

Zeng, G. et al. Adaptable and precise: Enterprise-scenario LLM function-calling capability training pipeline. arXiv preprint arXiv:2412.15660, 2024.
Publicado
19/10/2026
NOGUEIRA, João Conrado V.. STAB: A Complexity-Ladder Framework for Evaluating Local LLMs as Autonomous Agents. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 225-232. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.29261.