Avaliação simétrica de detecção e over-blocking na geração de políticas OPA/Rego por LLMs
Resumo
As métricas usadas para avaliar geração de políticas policy-as-code por LLMs medem detecção, mas quase nunca especificidade. Propomos a strict pass@1, métrica pareada construída sobre avaliação por plano dual: a política precisa disparar no plano Terraform vulnerável e ficar silenciosa no plano seguro espelhado (minimal pair). Em 30 cenários, 5 LLMs e 900 execuções (K = 3), comparamos Zero-shot (ZS) e Recursive Criticism and Improvement (RCI) com Compiler-in-the-Loop. O RCI aumenta a aprovação na verificação estática e a detecção, mas piora a strict pass@1 em quatro dos cinco modelos. No gpt-4o a queda vai de 28,9% a 0,0% por execução, sem que nenhum dos 30 cenários produzisse política utilizável sob RCI (p = 0,004, teste pareado por cenário). Medir apenas detecção, portanto, não caracteriza a utilidade operacional das políticas geradas.Referências
Bruni, M., Gabrielli, F., Ghafari, M., and Kropp, M. (2025). Benchmarking prompt engineering techniques for secure code generation with GPT models. arXiv preprint arXiv:2502.06039.
Chowdhary, N., Dutta, T., Chattopadhyay, S., and Chakraborty, S. (2025). AutoPAC: Exploring LLMs for automating policy to code conversion in business organizations. In 2025 17th International Conference on COMmunication Systems and NETworks (COMSNETS). IEEE.
Khan, R. N. H., Wasif, D., Cho, J.-H., and Butt, A. (2025). Multi-agent code-orchestrated generation for reliable infrastructure-as-code. arXiv preprint arXiv:2510.03902v1.
Kon, P. T. J., Liu, J., Qiu, Y., Fan, W., He, T., Lin, L., Zhang, H., Park, O. M., Elengikal, G. S., Kang, Y., Chen, A., Chowdhury, M., Lee, M., and Wang, X. (2024). IaC-Eval: A code generation benchmark for cloud infrastructure-as-code programs. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track.
Li, Y., Grella, M., Nahmias, D., Engelberg, G., Klein, D., Guizzardi, G., van Ede, T., and Continella, A. (2025). GENSIAC: Toward security-aware infrastructure-as-code generation with large language models. arXiv preprint arXiv:2511.12385.
Panickssery, A., Bowman, S. R., and Feng, S. (2024). LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076.
Romeo, F., Arena, L., Blefari, F., Pironti, F. A., Lupinacci, M., and Furfaro, A. (2025). ARPACCINO: An agentic-RAG for policy as code compliance. arXiv preprint arXiv:2507.10584v2.
Tony, C., Díaz Ferreyra, N. E., Mutas, M., Dhif, S., and Scandariato, R. (2025). Prompting techniques for secure code generation: A systematic investigation. arXiv preprint arXiv:2407.07064.
Zhang, T., Pan, S., Zhang, Z., Xing, Z., and Sun, X. (2025). Deployability-centric infrastructure-as-code generation: An LLM-based iterative framework. arXiv preprint arXiv:2506.05623.
Chowdhary, N., Dutta, T., Chattopadhyay, S., and Chakraborty, S. (2025). AutoPAC: Exploring LLMs for automating policy to code conversion in business organizations. In 2025 17th International Conference on COMmunication Systems and NETworks (COMSNETS). IEEE.
Khan, R. N. H., Wasif, D., Cho, J.-H., and Butt, A. (2025). Multi-agent code-orchestrated generation for reliable infrastructure-as-code. arXiv preprint arXiv:2510.03902v1.
Kon, P. T. J., Liu, J., Qiu, Y., Fan, W., He, T., Lin, L., Zhang, H., Park, O. M., Elengikal, G. S., Kang, Y., Chen, A., Chowdhury, M., Lee, M., and Wang, X. (2024). IaC-Eval: A code generation benchmark for cloud infrastructure-as-code programs. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track.
Li, Y., Grella, M., Nahmias, D., Engelberg, G., Klein, D., Guizzardi, G., van Ede, T., and Continella, A. (2025). GENSIAC: Toward security-aware infrastructure-as-code generation with large language models. arXiv preprint arXiv:2511.12385.
Panickssery, A., Bowman, S. R., and Feng, S. (2024). LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076.
Romeo, F., Arena, L., Blefari, F., Pironti, F. A., Lupinacci, M., and Furfaro, A. (2025). ARPACCINO: An agentic-RAG for policy as code compliance. arXiv preprint arXiv:2507.10584v2.
Tony, C., Díaz Ferreyra, N. E., Mutas, M., Dhif, S., and Scandariato, R. (2025). Prompting techniques for secure code generation: A systematic investigation. arXiv preprint arXiv:2407.07064.
Zhang, T., Pan, S., Zhang, Z., Xing, Z., and Sun, X. (2025). Deployability-centric infrastructure-as-code generation: An LLM-based iterative framework. arXiv preprint arXiv:2506.05623.
Publicado
01/09/2026
Como Citar
RODRIGUES, Eduardo; SANTIN, Altair Olivo; VIEGAS, Eduardo K.; VEIGA, Fellipe Medeiros.
Avaliação simétrica de detecção e over-blocking na geração de políticas OPA/Rego por LLMs. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 404-418.
DOI: https://doi.org/10.5753/sbseg_estendido.2026.29495.
