Defense-in-Depth for Generative AI: An Architectural Evaluation and Comparison of LLM Guardrails
Resumo
This paper analyzes multi-layered guardrail architectures designed to secure LLMs at runtime. By systematically comparing five frameworks (Prompt Guard, Llama Guard, Llama Firewall, NeMo, and Amazon Bedrock), the study evaluates their structural designs and operational trade-offs through a qualitative architectural systematization. Ultimately, it concludes that selecting the appropriate architecture requires a systematic approach to balance security rigor and model utility.
Referências
Amazon Web Services (2025) “Guardrails for Amazon Bedrock”, available at: [link].
Azambuja, A. J. et al. (2026) “Evaluation of NeMo Guardrails as a Firewall for User–LLM Interaction”, Future Internet, 18(5), article number 252, DOI: 10.3390/fi18050252.
Banwasi, A., Friedman, S. M. and Khanzadeh, M. (2024) “A Comparative Analysis of Guardrail Frameworks for Large Language Models and Enhancement with Ensemble Techniques”, In: New York: Columbia Univ. School of Eng. and Applied Science.
Chao, P. et al. (2024) "JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models", In: Proceedings of the 38th Conference on NeurIPS 2024, doi: [link]
Chennabasappa, S. et al. (2025) “LlamaFirewall: an open source guardrail system for building secure AI agents”, Meta AI, available at: [link].
Devino, M., Ju, E. and Caldeira Junior, P. M. (2025) “Designing and implementing LLM guardrails components in production environments”, In: Proceedings of the 2025 IEEE/ACM International Conference on AI Engineering – Software Engineering for AI, pages 12–17, DOI: 10.1109/CAIN66642.2025.00010.
Dong, Y. et al. (2024) “Position: building guardrails for large language models requires systematic design”, In: Proc. of the 41st ICML'24, article number 451, 20 pages, JMLR.org, DOI: 10.5555/3692070.3692521.
Dong, Y. et al. (2025) “Safeguarding large language models: a survey”, Artificial Intelligence Review, 58(12), article number 382, DOI: 10.1007/s10462-025-11389-2.
Inan, H. et al. (2023) “Llama Guard: LLM-based input-output safeguard for human-AI conversations”, Meta AI, available at: [link].
Leon, J. et al. (2024) "Garak: A Framework for Security Probing Large Language Models", DOI: 10.48550/arXiv.2406.11036
Malik, V. et al. (2025) “AI-Native LLM Security: Threats, Defenses, and Best Practices for Building Safe and Trustworthy AI”, Packt Publishing, Birmingham.
Mazeika, M. et al. (2024) "HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Alignment", In: Proc. of the 41st ICML'24, doi: [link].
MITRE ATLAS (2025) “Generative AI Guardrails - Mitigation AML.M0020”, available at: [link].
NIST (2023) "Artificial Intelligence Risk Management Framework (AI RMF 1.0)", NIST Trustworthy and Responsible AI, NIST AI 100-1.
NIST (2025) “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations”, NIST Trustworthy and Responsible AI Report, NIST AI 100-2, DOI: 10.6028/NIST.AI.100-2.
OWASP (2025) “OWASP Top 10 for Large Language Model Applications”, OWASP Foundation.
OWASP (2026) “AI Security Overview - The AI Exchange Framework”, OWASP Foundation..
Paim, K. O. et al. (2025) “Exploiting latent space discontinuities for building universal LLM jailbreaks and data extraction attacks”, In: Anais do XXV SBSeg, pages 417–431, DOI: 10.5753/sbseg.2025.11448.
Rebedea, T. et al. (2023) “NeMo Guardrails: a toolkit for controllable and safe LLM applications with programmable rails”, In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Association for Computational Linguistics, Singapore, pages 431–445, DOI: 10.18653/v1/2023.emnlp-demo.40.
Russinovich, M., Salem, A. and Ronen, R. (2024) "Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack", Microsoft, DOI: 10.48550/arXiv.2404.01833.
Tripathi, P. et al. (2026) “Guardrail-based approaches for enhancing safety, security, and privacy in large language models”, In: Proc. of the 2026 IC3ECSBHI, pages 1619–1624, DOI: 10.1109/IC3ECSBHI67834.2026.11469093.
