A Synthetic Benchmark for Integrated Security Assessment of RAG-Based Corporate Assistants
Resumo
This paper presents a synthetic benchmark for integrated security assessment of RAG-based corporate assistants. The benchmark models a fictitious organization and instruments security-relevant stages across the RAG pipeline. It comprises 473 cases per configuration, including 352 adversarial cases organized into 23 attack families, and evaluates ten defensive configurations. The analysis combines family-macro Attack Success Rate with stage-level metrics to distinguish internal pipeline violations from failures observable in the final response. Results show that output-level mitigations may suppress visible leakage while leaving upstream authorization and retrieval failures unresolved, whereas layered observable defenses provide broader protection while preserving benign utility. An exploratory comparison with a local Llama model shows the same directional defensive pattern under a different generation mechanism. These findings support pipeline-level evaluation as a more informative approach to assessing RAG-based corporate assistants than final-response analysis alone.
Referências
Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., and Wong, E. (2024). JailbreakBench: An open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems, volume 37.
Clop, C. and Teglia, Y. (2024). Backdoored retrievers for prompt injection attacks on retrieval-augmented generation of large language models. arXiv preprint arXiv:2410.14479. [link].
Dubey, A. et al. (2024). The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. [link].
Figueira, L. d. O., Duarte, L. O., and Silva, R. L. (2026). Integrated security assessment of rag-based corporate assistants. Open-science artifact package. [link]. Accessed August 12, 2026.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. arXiv preprint arXiv:2302.12173. [link].
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474.
Liang, X., Niu, S., Li, Z., Zhang, S., Wang, H., Xiong, F., Fan, Z., Tang, B., Zhao, J., Yang, J., Song, S., and Wang, M. (2025). SafeRAG: Benchmarking security in retrieval-augmented generation of large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4609–4631, Vienna, Austria. Association for Computational Linguistics. DOI: 10.18653/v1/2025.acl-long.230.
Liu, X. et al. (2024). Automatic and universal prompt injection attacks against large language models. arXiv preprint arXiv:2403.04957. [link] 2403.04957.
Open Worldwide Application Security Project (2024). OWASP top 10 for LLM applications 2025. Technical report, Open Worldwide Application Security Project. [link].
Open Worldwide Application Security Project (2026). OWASP top 10 for LLM applications 2026. Technical report, Open Worldwide Application Security Project. [link].
Silva, R. L. (2026). Toward quantum-resilient software supply chains: A DevSecOps case study with hybrid post-quantum artifact signing. In Proceedings of the 2026 IEEE International Conference on Quantum Communications, Networking, and Computing (QCNC), Kobe, Japan. IEEE. Accepted paper.
Silva, R. L. and Duarte, L. O. (2026a). Engineering trust in LLM supply chains through hybrid post-quantum artifact signatures. In Proceedings of the 20th IEEE International Workshop on Security, Trust, and Privacy for Software Applications (STPSA 2026), IEEE COMPSAC 2026, Madrid, Spain. IEEE. Presented at STPSA 2026.
Silva, R. L. and Duarte, L. O. (2026b). Securing LLM software supply chains: A layered lifecycle framework with hybrid post-quantum artifact signing. In Proceedings of IEEE IDS 2026, New York City, USA. IEEE. Presented at IEEE IDS 2026.
Wei, A., Haghtalab, N., and Steinhardt, J. (2023). Jailbroken: How does LLM safety training fail? arXiv preprint arXiv:2307.02483. [link].
Yi, J., Xie, Y., Zhu, B., Kiciman, E., Sun, G., Xie, X., and Wu, F. (2025). Benchmarking and defending against indirect prompt injection attacks on large language models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery. DOI: 10.1145/3690624.3709179.
Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Gong, N. Z., Zhang, Y., and Xie, X. (2024). PromptRobust: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528. [link].
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043. [link].
