AIaCGateGuard: A Security-First Pipeline for Benchmarking LLM- and SLM-Generated Infrastructure-as-Code

Resumo


We present AIACGATEGUARD, an open-source, fully automated pipeline for evaluating the security compliance of Infrastructure-as-Code (IaC) generated by Large and Small Language Models (LLMs/SLMs). Unlike existing IaC generation benchmarks, which assess only syntactic validity or functional correctness, AIACGATEGUARD integrates dual-tool static security analysis (Checkov and Trivy) directly into a GitLab CI/CD pipeline, producing permodel, per-scenario security compliance reports. The tool supports plug-and-play integration of new LLMs via REST API and new SLMs via local inference (llama-server), configurable prompt strategies and security specificity levels, optional infrastructure validation and deployment stages (plan-real, apply, destroy), and a Model Context Protocol (MCP) server for conversational interaction with the pipeline. We demonstrate the tool on seven state-of-the-art models across 17 AWS Terraform scenarios, producing 3,570 automated evaluations and exposing security compliance gaps invisible to syntactic-only validation. AIACGATEGUARD is publicly available, including pipeline code, prompt templates, scenario definitions, and a video demonstration of installation and usage.

Referências

Anthropic (2024). Model context protocol. [link].

Aqua Security (2024). Trivy: Comprehensive security scanner. [link].

Bridgecrew (2024). Checkov: Static code analysis for infrastructure-as-code. [link].

Brown, S. (2018). Software Architecture for Developers. Leanpub. C4 model: [link].

Kon, P. T. J., Liu, J., Qiu, Y., Fan, W., He, T., Lin, L., Zhang, H., Park, O. M., Elengikal, G. S., Kang, Y., et al. (2024). IaC-Eval: A code generation benchmark for cloud infrastructure-as-code programs. In Advances in Neural Information Processing Systems (NeurIPS), volume 37, pages 134488–134512.

Morris, K. (2020). Infrastructure as Code: Dynamic Systems for the Cloud Age. O’Reilly Media, 2nd edition.

Vargas, F. L. S., Mansilha, R. B., and Kreutz, D. (2026a). Can language models generate secure Terraform code? A security-focused benchmark using static analysis. In Anais do I Simpósio de Infraestrutura Digital/Nuvem para Pesquisa (Pesquisa@Nuvem), pages 29–38. SBC.

Vargas, F. L. S., Mansilha, R. B., and Kreutz, D. (2026b). Security-first evaluation of text-to-Terraform: Benchmarking LLMs and SLMs for secure IaC generation. In Anais do XXVI Simpósio Brasileiro de Segurança da Informação (SBSeg 2026). To be published.

Zhang, T., Pan, S., Zhang, Z., Xing, Z., and Sun, X. (2025). Deployability-centric infrastructure-as-code generation: An llm-based iterative framework. arXiv preprint arXiv:2506.05623.
Publicado
01/09/2026
VARGAS, Francis Luis Santos; MANSILHA, Rodrigo Brandão; KREUTZ, Diego. AIaCGateGuard: A Security-First Pipeline for Benchmarking LLM- and SLM-Generated Infrastructure-as-Code. In: SALÃO DE FERRAMENTAS - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 91-99. DOI: https://doi.org/10.5753/sbseg_estendido.2026.33594.

Artigos mais lidos do(s) mesmo(s) autor(es)

<< < 7 8 9 10 11 12