Who Guards the Guard? Evaluating Deterministic, LLM-Based, and Hybrid Firewalls for Tool-Using Agents

  • Camilla B. Quincozes UFU
  • Paulo Souza UNIPAMPA
  • Rafael Araujo UFU
  • Diego Molinos UFU
  • Silvio E. Quincozes UNIPAMPA

Resumo


Tool-using large language model (LLM) agents process untrusted files, URLs, and tables at a boundary where data may be mistaken for instructions. This paper compares four interface-level conditions: no firewall, deterministic policy enforcement, an LLM-based Tool-Input Minimizer/Tool-Output Sanitizer, and a gated hybrid. The evaluation has two stages. The development microbenchmark contains 19 adversarial and 4 benign cases, 6 Groq-hosted guard backbones, 2 repetitions, and 1,104 executions. A second frozen-policy held-out study contains 42 unseen adversarial variants and 20 benign hard negatives. Its primary analysis uses four complete guard backbones, three repetitions for LLM-dependent conditions, and case-level aggregation. In the development set, the hybrid condition exhibited no successful attacks among 228 adversarial executions. This result did not generalize: on the held-out set, both the deterministic and hybrid conditions achieved a case-mean attack success rate (ASR) of 54.76%, because the hybrid router invoked the semantic guard in 0/744 executions. The LLMonly condition reached 50.60% case-mean ASR, 92.08% benign utility, and 33/744 structured-output failures under a fail-open policy. The results show that deterministic checks are suitable for explicit structural constraints, but composition alone is insufficient; hybrid security depends on routing coverage, guard availability, and evaluation against attacks not used in constructing the defense.

Referências

Bhagwatkar, R., Kasa, K., Puri, A., Huang, G., Rish, I., Taylor, G. W., Dvijotham, K. D., and Lacoste, A. (2025). Indirect prompt injections: Are firewalls all you need, or stronger benchmarks? Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In The Thirty-Eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track.

Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security. ACM.

Liu, Y., Deng, G., Li, Y., Wang, K., Wang, Z., Wang, X., Zhang, T., Liu, Y., Wang, H., Zheng, Y., and Liu, Y. (2023). Prompt injection attack against LLM-integrated applications.

OWASP, G. S. P. (2025). OWASP top 10 for large language model applications 2025. [link]. Accessed: 2026-06-04.

Podpora, M., Baranowski, M., Chopcian, M., Kwasniewicz, L., and Radziewicz, W. (2026). Llm firewall using validator agent for prevention against prompt injection attacks. Applied Sciences, 16(1).

Shi, T., Zhu, K., Wang, Z., Jia, Y., Cai, W., Liang, W., Wang, H., Alzahrani, H., Lu, J., Kawaguchi, K., Alomair, B., Zhao, X., Wang, W. Y., Gong, N., Guo, W., and Song, D. (2025). PromptArmor: Simple yet effective prompt injection defenses.

Weng, S., Feng, Y., Zhang, J., Xie, X., Yu, J., and Liu, J. (2026). ARGUS: Defending LLM agents against context-aware prompt injection.

Yao, S., Shinn, N., Razavi, P., and Narasimhan, K. (2024). τ -Bench: A benchmark for tool-agent-user interaction in real-world domains.

Zhan, Q., Liang, Z., Ying, Z., and Kang, D. (2024). InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents. In Findings of the Association for Computational Linguistics: ACL 2024, pages 10471–10506, Bangkok, Thailand. Association for Computational Linguistics.

Zhang, H., Huang, J., Mei, K., Yao, Y., Wang, Z., Zhan, C., Wang, H., and Zhang, Y. (2024). Agent security bench (ASB): Formalizing and benchmarking attacks and defenses in LLM-based agents.

Zhao, W., Li, Z., Zhang, P., and Sun, J. (2026). ClawGuard: A runtime security framework for tool-augmented LLM agents against indirect prompt injection.
Publicado
01/09/2026
QUINCOZES, Camilla B.; SOUZA, Paulo; ARAUJO, Rafael; MOLINOS, Diego; QUINCOZES, Silvio E.. Who Guards the Guard? Evaluating Deterministic, LLM-Based, and Hybrid Firewalls for Tool-Using Agents. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 774-784. DOI: https://doi.org/10.5753/sbseg_estendido.2026.29812.