UnifiedBench: Toward a Unified Dataset for Red-Teaming Language Models

  • Matheus Azevedo de Sá UDESC
  • Paulo Henrique Gomes de Senna UDESC
  • Milton Pedro Pagliuso Neto UDESC
  • Charles Christian Miers UDESC

Resumo


Adversarial prompting constitutes one of the principal threats to deployed Large Language Models, yet the evaluation of alignment robustness against such attacks is hindered by the fragmentation of existing benchmark datasets. AdvBench, HarmBench, and JailbreakBench — the three most adopted adversarial prompt datasets — adopt mutually incompatible categorical schemes, which prevents cross-dataset comparison of category-level vulnerability. We address this gap through four contributions: (i) a unified taxonomy of harmful behaviors anchored in the NIST AI 100-2 framework, comprising seven categories that decompose the Misuse vulnerability class; (ii) UnifiedBench, a consolidated dataset of 810 adversarial prompts derived from the three source datasets and categorized under this taxonomy; (iii) a query-only evaluation pipeline combining an automated judge (LlamaGuard-3-1B) with human reviewers to measure per-category Attack Success Rate (ASR); and (iv) a comparative analysis of two instruction-tuned models under direct harmful prompting. Our findings reveal current alignment procedures are comparatively less effective at preserving deployment-specific operational constraints than at preventing conventionally harmful content.

Referências

Beyer, L. et al. (2025). Llm-safety evaluations lack robustness. In International Conference on Machine Learning (ICML). Peer-reviewed.

Busch, F., Hoffmann, L., dos Santos, D. P., et al. (2025). Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine, 5.

Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramer, F., et al. (2024a). Jailbreakbench: An open robustness benchmark for jailbreaking large language models. Advances in Neural Information Processing Systems, 37:55005–55029.

Chao, P., Debenedetti, E., Robey, A., et al. (2024b). Jailbreakbench: An open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems (NeurIPS). Peer-reviewed.

Huang, Y. et al. (2023). Maliciousinstruct: A large-scale benchmark for safety alignment of large language models. arXiv preprint arXiv:2310.06987.

Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M. (2023). Llama Guard: LLM-based input-output safeguard for Human-AI conversations. arXiv preprint arXiv:2312.06674.

Ji, J. et al. (2023). Beavertails: Towards improved safety alignment of llm via a human-preference dataset. In Advances in Neural Information Processing Systems (NeurIPS). Peer-reviewed.

Khan, J. A. et al. (2026). Recommendations for efficient and responsible LLM adoption within industrial software development. Journal of Systems and Software.

Landis, J. R. and Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1):159–174.

Li, M. Q. and Fung, B. C. M. (2025). Security concerns for large language models: A survey. arXiv preprint arXiv:2505.18889.

Li, Y. et al. (2024). Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. In Findings of the Association for Computational Linguistics (ACL Findings). Peer-reviewed.

Lin, Z. et al. (2023). Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversations. In Findings of the Association for Computational Linguistics (EMNLP Findings). Peer-reviewed.

Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., et al. (2024a). Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. In International Conference on Machine Learning (ICML). Peer-reviewed.

Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. (2024b). Harmbench: a standardized evaluation framework for automated red teaming and robust refusal. In ICML.

Menlo Ventures (2025). The state of generative AI in the enterprise 2025. Accessed: 2026-05-11.

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS).

OWASP Foundation (2025). OWASP top 10 for LLM applications 2025. Accessed: 2026-05-11.

Röttger, P. et al. (2024). Xstest: A test suite for identifying excessive refusals in large language models. In Annual Conference of the North American Chapter of the ACL (NAACL). Peer-reviewed.

Shayegani, E., Mamun, M. A. A., Fu, Y., Zaree, P., Dong, Y., and Abu-Ghazaleh, N. (2023). Survey of vulnerabilities in large language models revealed by adversarial attacks. arXiv preprint arXiv:2310.10844.

Sun, H. et al. (2024). Trustllm: Trustworthiness in large language models. In International Conference on Machine Learning (ICML). Peer-reviewed.

Vassilev, A., Oprea, A., Fordyce, A., Anderson, H., Davies, X., and Hamin, M. (2025). Adversarial machine learning: A taxonomy and terminology of attacks and mitigations. Technical report, NIST AI 100-2e2025.

Vidgen, B. et al. (2023). Simplesafetytests: A test suite for safety evaluation of language models. arXiv preprint arXiv:2308.08469. Preprint widely adopted by the LLM safety community.

Wang, B. et al. (2023). Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. In Advances in Neural Information Processing Systems (NeurIPS). Peer-reviewed.

Wang, Y. et al. (2024). Do-not-answer: A dataset for evaluating safeguards in llms. In Findings of the European Chapter of the ACL (EACL Findings). Peer-reviewed.

Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., et al. (2022). Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT), pages 214–229.

Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., et al. (2025). Qwen3 technical report. Preprint arXiv:2505.09388.

Zhang, X. et al. (2024). Safetybench: Evaluating the safety of large language models. In Annual Meeting of the Association for Computational Linguistics (ACL). Peer-reviewed.

Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems (NeurIPS).

Zhu, W., Zhao, Z., Chen, Y., Wang, Y., and Xie, X. (2023). Promptbench: A unified library for evaluation of large language models. arXiv preprint arXiv:2312.07910.

Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M. (2023a). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M. (2023b). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.
Publicado
01/09/2026
SÁ, Matheus Azevedo de; SENNA, Paulo Henrique Gomes de; PAGLIUSO NETO, Milton Pedro; MIERS, Charles Christian. UnifiedBench: Toward a Unified Dataset for Red-Teaming Language Models. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 762-773. DOI: https://doi.org/10.5753/sbseg_estendido.2026.29187.

Artigos mais lidos do(s) mesmo(s) autor(es)