PAIR transfers, CipherChat does not: Indirect Prompt Injection in MCP Tool Responses

  • Matheus R. S. Corrêa UDESC
  • Carlos D. S. Bunn UDESC
  • Charles C. Miers UDESC
  • Milton P. P. Neto UDESC

Resumo


The Model Context Protocol (MCP) standardizes the integration between Large Language Models (LLMs) and external tools, enabling agents to act on real environments and expanding the attack surface beyond conventional text interfaces. This paper evaluates two black-box attack strategies against an MCP based agent. Prompt Automatic Iterative Refinement (PAIR), which iteratively refines injected instructions via attacker model feedback, and CipherChat, adapted to the indirect injection context of MCP tool responses. Attack success is redefined as the invocation of an unauthorized tool action rather than harmful text output, reflecting the qualitatively different threat posed by agentic systems. The results show that semantic adversarial refinement transfers to the MCP tool-response channel and that larger models offer partial but insufficient protection, whereas encoding based obfuscation fails to transfer from the direct user-turn context to the MCP data plane.

Referências

Anthropic (2024). Introducing the model context protocol. [link]. Accessed: 2026-03-20.

Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. (2024). Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419.

Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., and Vechev, M. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. arXiv preprint arXiv:2406.13352.

Fischer, B.-K. (2025). MCP security notification: Tool poisoning attacks. [link].

Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. arXiv preprint arXiv:2302.12173.

Huang, C. et al. (2026). Model context protocol threat modeling and analyzing vulnerabilities to prompt injection with tool poisoning. arXiv preprint arXiv:2603.22489.

MITRE (2026). AI security 101 | MITRE ATLAS. [link].

OWASP (2026). OWASP top 10 for large language model applications. [link].

Perez, F. and Ribeiro, I. (2022). Ignore previous prompt: Attack techniques for language models. [link].

Radosevich, B. and Halloran, J. (2025). MCP safety audit: LLMs with the model context protocol allow major security exploits. arXiv preprint arXiv:2504.03767.

Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., and Hashimoto, T. (2023). Identifying the risks of LM agents with an LM-emulated sandbox. arXiv preprint arXiv:2309.15817.

Wang, Z. et al. (2025). MCPTox: A benchmark for tool poisoning attack on real-world MCP servers. arXiv preprint arXiv:2508.14925.

Wei, A., Haghtalab, N., and Steinhardt, J. (2023). Jailbroken: How does LLM safety training fail? arXiv preprint arXiv:2307.02483.

Yuan, Y., Jiao, W., Wang, W., Huang, J.-t., He, P., Shi, S., and Tu, Z. (2023). GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher. arXiv preprint arXiv:2308.06463.

Zou, A., Wang, Z., Carlini, N., Nasr, Miladand Kolter, J. Z., and Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.
Publicado
01/09/2026
CORRÊA, Matheus R. S.; BUNN, Carlos D. S.; MIERS, Charles C.; P. NETO, Milton P.. PAIR transfers, CipherChat does not: Indirect Prompt Injection in MCP Tool Responses. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 638-648. DOI: https://doi.org/10.5753/sbseg_estendido.2026.29477.

Artigos mais lidos do(s) mesmo(s) autor(es)

1 2 > >>