Data Exfiltration in Model Context Protocol (MCP)-Based Intelligent Agents: An Evaluation of Prompt Injection Vectors
Resumo
The integration of Intelligent Agents utilizing Large Language Models (LLMs) with external systems via the Model Context Protocol (MCP) expands the capabilities of autonomous agents, but also introduces critical data exfiltration risks. This study empirically evaluates the native defenses of Intelligent Agents against direct technical commands and social engineering-based prompt injections. In 160 automated tests across eight models, the results indicate that current semantic alignment is insufficient: models resistant to direct orders became vulnerable in seemingly authorized contexts, exhibiting up to 40% data leakage, while security-oriented models reached up to 90% successful exfiltration of sensitive files. We conclude that MCP architectures should not rely solely on the safety barriers of the underlying model, requiring strict egress controls at the protocol layer.Referências
Agno (2026). Agno: Build multi-modal agents with memory, knowledge and tools. [link]. Acessado em: 06 de maio de 2026.
Alibaba Group (2026). Qwen 3 32b: Dense large language model technical overview. Acessado em: 9 de maio de 2026.
Anthropic (2024). Introducing the model context protocol. Accessed: 2026-05-19.
Correia, P. H. B., Achjian, R. W., de Oliveira, D. E. G. C., Maria, Y. A., Hayashi, V. T., Lopes, M., Miers, C. C., and Jr, M. A. S. (2026). A systematic literature review on llm defenses against prompt injection and jailbreaking: Expanding nist taxonomy.
Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In Advances in Neural Information Processing Systems, volume 37, pages 82895–82920.
DeepSeek (2026). Deepseek api docs. [link]. Acessado em: 06 de maio de 2026.
Ferrag, M. A., Lakas, A., Tihanyi, N., and Debbah, M. (2025). Securing LLM agents: From prompt sanitization to autonomous red teaming and beyond. Internet of Things and Cyber-Physical Systems, 5:185–209.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec ’23, pages 79–90. Association for Computing Machinery.
Groq Inc. (2026). Groq cloud: Fast ai inference. [link]. Acessado em: 06 de maio de 2026.
Guo, Y., Liu, P., Ma, W., Deng, Z., Zhu, X., Di, P., Xiao, X., and Wen, S. (2025). Systematic analysis of mcp security.
Hou, X., Zhao, Y., Wang, S., and Wang, H. (2026). Model context protocol (MCP): Landscape, security threats, and future research directions. ACM Transactions on Software Engineering and Methodology. Online ahead of print; arXiv:2503.23278.
Huang, C., Huang, X., Tran, N. P., and Milani Fard, A. (2026). Model context protocol threat modeling and analysis of vulnerabilities to prompt injection with tool poisoning. Journal of Cybersecurity and Privacy, 6(3):84.
Kumar, S. S., Cummings, M. L., and Stimpson, A. (2024). Strengthening llm trust boundaries: A survey of prompt injection attacks.
Maloyan, N. and Namiot, D. (2026). Breaking the protocol: Security analysis of the model context protocol specification and prompt injection vulnerabilities in tool-integrated llm agents.
Meta AI (2025). Llama 3.3: 70b versatile model card and specifications. Acessado em: 9 de maio de 2026.
Model Context Protocol Contributors (2025). Model Context Protocol specification. [link]. Versão 2025-11-25. Acesso em: 08 mai. 2026.
OpenAI (2025). Technical report: Performance and baseline evaluations of gpt-osssafeguard-120b and gpt-oss-safeguard-20b. Technical report, OpenAI. Acessado em: 9 de maio de 2026.
OpenRouter (2026). Openrouter: A unified interface for large language models. [link]. Acessado em: 06 de maio de 2026.
OWASP Foundation (2025). LLM01:2025 prompt injection. [link]. OWASP Top 10 for Large Language Model Applications; Accessed: 2026-05-19.
Poolside AI (2026). Laguna m.1: 225b mixture-of-experts for agentic coding. Acessado em: 9 de maio de 2026.
Python Software Foundation (2026). Python language reference, version 3.x. [link]. Acessado em: 06 de maio de 2026.
Rall, D., Bauer, B., Mittal, M., and Fraunholz, T. (2025). Exploiting web search tools of ai agents for data exfiltration.
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
Ullah, F., Edwards, M., Ramdhany, R., Chitchyan, R., Babar, M. A., and Rashid, A. (2018). Data exfiltration: A review of external attack vectors and countermeasures. Journal of Network and Computer Applications, 101:18–54.
Valerian, D., Winoto, I., Ma, Y., Tsukamoto, K., and Shao, C. (2026). Intelligent chatbot system that integrates large language models with the model context protocol. In Barolli, L., Miwa, H., and Natwichai, J., editors, Advances in Intelligent Networking and Collaborative Systems, pages 252–259, Cham. Springer Nature Switzerland.
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., and Wen, J. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):186345.
Wang, N., Walter, K., Gao, Y., and Abuadbba, A. (2026). Understanding the adversarial landscape of large language models through the lens of attack objectives. IEEE Security & Privacy, 24(1):53–60.
Wei, A., Haghtalab, N., and Steinhardt, J. (2023). Jailbroken: How does llm safety training fail? Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Qin, W., Zheng, Y., Qiu, X., Huang, X., Zhang, Q., and Gui, T. (2025). The rise and potential of large language model based agents: a survey. Science China Information Sciences, 68(2):121101.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022). React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
Zhan, Q., Liang, Z., Ying, Z., and Kang, D. (2024). InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 10471–10506, Bangkok, Thailand. Association for Computational Linguistics.
Zhipu AI (2026). Glm-4.5: Hybrid reasoning model technical report. Acessado em: 9 de maio de 2026.
Alibaba Group (2026). Qwen 3 32b: Dense large language model technical overview. Acessado em: 9 de maio de 2026.
Anthropic (2024). Introducing the model context protocol. Accessed: 2026-05-19.
Correia, P. H. B., Achjian, R. W., de Oliveira, D. E. G. C., Maria, Y. A., Hayashi, V. T., Lopes, M., Miers, C. C., and Jr, M. A. S. (2026). A systematic literature review on llm defenses against prompt injection and jailbreaking: Expanding nist taxonomy.
Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In Advances in Neural Information Processing Systems, volume 37, pages 82895–82920.
DeepSeek (2026). Deepseek api docs. [link]. Acessado em: 06 de maio de 2026.
Ferrag, M. A., Lakas, A., Tihanyi, N., and Debbah, M. (2025). Securing LLM agents: From prompt sanitization to autonomous red teaming and beyond. Internet of Things and Cyber-Physical Systems, 5:185–209.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec ’23, pages 79–90. Association for Computing Machinery.
Groq Inc. (2026). Groq cloud: Fast ai inference. [link]. Acessado em: 06 de maio de 2026.
Guo, Y., Liu, P., Ma, W., Deng, Z., Zhu, X., Di, P., Xiao, X., and Wen, S. (2025). Systematic analysis of mcp security.
Hou, X., Zhao, Y., Wang, S., and Wang, H. (2026). Model context protocol (MCP): Landscape, security threats, and future research directions. ACM Transactions on Software Engineering and Methodology. Online ahead of print; arXiv:2503.23278.
Huang, C., Huang, X., Tran, N. P., and Milani Fard, A. (2026). Model context protocol threat modeling and analysis of vulnerabilities to prompt injection with tool poisoning. Journal of Cybersecurity and Privacy, 6(3):84.
Kumar, S. S., Cummings, M. L., and Stimpson, A. (2024). Strengthening llm trust boundaries: A survey of prompt injection attacks.
Maloyan, N. and Namiot, D. (2026). Breaking the protocol: Security analysis of the model context protocol specification and prompt injection vulnerabilities in tool-integrated llm agents.
Meta AI (2025). Llama 3.3: 70b versatile model card and specifications. Acessado em: 9 de maio de 2026.
Model Context Protocol Contributors (2025). Model Context Protocol specification. [link]. Versão 2025-11-25. Acesso em: 08 mai. 2026.
OpenAI (2025). Technical report: Performance and baseline evaluations of gpt-osssafeguard-120b and gpt-oss-safeguard-20b. Technical report, OpenAI. Acessado em: 9 de maio de 2026.
OpenRouter (2026). Openrouter: A unified interface for large language models. [link]. Acessado em: 06 de maio de 2026.
OWASP Foundation (2025). LLM01:2025 prompt injection. [link]. OWASP Top 10 for Large Language Model Applications; Accessed: 2026-05-19.
Poolside AI (2026). Laguna m.1: 225b mixture-of-experts for agentic coding. Acessado em: 9 de maio de 2026.
Python Software Foundation (2026). Python language reference, version 3.x. [link]. Acessado em: 06 de maio de 2026.
Rall, D., Bauer, B., Mittal, M., and Fraunholz, T. (2025). Exploiting web search tools of ai agents for data exfiltration.
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
Ullah, F., Edwards, M., Ramdhany, R., Chitchyan, R., Babar, M. A., and Rashid, A. (2018). Data exfiltration: A review of external attack vectors and countermeasures. Journal of Network and Computer Applications, 101:18–54.
Valerian, D., Winoto, I., Ma, Y., Tsukamoto, K., and Shao, C. (2026). Intelligent chatbot system that integrates large language models with the model context protocol. In Barolli, L., Miwa, H., and Natwichai, J., editors, Advances in Intelligent Networking and Collaborative Systems, pages 252–259, Cham. Springer Nature Switzerland.
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., and Wen, J. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):186345.
Wang, N., Walter, K., Gao, Y., and Abuadbba, A. (2026). Understanding the adversarial landscape of large language models through the lens of attack objectives. IEEE Security & Privacy, 24(1):53–60.
Wei, A., Haghtalab, N., and Steinhardt, J. (2023). Jailbroken: How does llm safety training fail? Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Qin, W., Zheng, Y., Qiu, X., Huang, X., Zhang, Q., and Gui, T. (2025). The rise and potential of large language model based agents: a survey. Science China Information Sciences, 68(2):121101.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022). React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
Zhan, Q., Liang, Z., Ying, Z., and Kang, D. (2024). InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 10471–10506, Bangkok, Thailand. Association for Computational Linguistics.
Zhipu AI (2026). Glm-4.5: Hybrid reasoning model technical report. Acessado em: 9 de maio de 2026.
Publicado
01/09/2026
Como Citar
BARCELOS, Tuigg R.; QUINCOZES, Silvio E.; SOUZA, Paulo.
Data Exfiltration in Model Context Protocol (MCP)-Based Intelligent Agents: An Evaluation of Prompt Injection Vectors. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 440-455.
DOI: https://doi.org/10.5753/sbseg.2026.28936.
