Jogos de Sinalização para Agentes de IA Comprometidos

  • Lucila M. S. Bento UERJ
  • Davidson R. Boccardo Einstein Hospital Israelita

Resumo


Agentes de IA podem possuir credenciais delegadas válidas e operar sob um contexto comprometido. Neste trabalho, a autorização adaptativa é modelada como um jogo de sinalização Bayesiano, no qual um agente íntegro ou comprometido emite um sinal básico ou reforçado, um detector produz evidência ruidosa e o verificador concede, desafia ou nega acesso. São derivados limiares e condições de separação, comparadas políticas binária e graduada e aplicado o refinamento D1 por programação linear. Em uma avaliação computacional com parâmetros sintéticos, a política graduada apresenta utilidade estritamente maior em uma região de crenças posteriores para 38,2% das configurações avaliadas. Entretanto, contenções mais fortes podem estabilizar pooling, mostrando que maior utilidade defensiva não implica separação dos tipos.

Referências

Banks, J. S. and Sobel, J. (1987). Equilibrium selection in signaling games. Econometrica, 55(3):647–661.

Boudagdigue, C., Benslimane, A., Kobbane, A., and Liu, J. (2023). Trust-based certificate management for industrial iot networks. IEEE Internet of Things Journal, 10(14):12867– 12885.

Cho, I.-K. and Kreps, D. M. (1987). Signaling games and stable equilibria. The Quarterly Journal of Economics, 102(2):179–221.

Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F. Agentdojo: a dynamic environment to evaluate prompt injection attacks and defenses for llm agents. In Advances in Neural Information Processing Systems 37, NIPS ’24.

Esposito, C., Tamburis, O., Su, X., and Choi, C. (2020). Robust decentralised trust management for the internet of things by using game theory. Information Processing & Management, 57(6):102308.

Fujie, N., Suzuki, S., Kurosaka, T., and Abe, R. (2026). Ieee infocom 2026 - ieee conference on computer communications. pages 1–7.

Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec ’23, page 79–90, New York, NY, USA. Association for Computing Machinery.

Jia, F., Wu, T., Qin, X., and Squicciarini, A. (2025). The task shield: Enforcing task alignment to defend against indirect prompt injection in LLM agents. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T., editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 29680–29697, Vienna, Austria. Association for Computational Linguistics.

Kolhe, A., Raju Dantuluri, V. S., Rajput, D., and Bajaj, S. (2026). Entitlement-aware authorization for mcp-based ai search and chat systems. In 2026 International Conference on Artificial Intelligence, Systems, and Emerging Technologies (ICAISET), pages 1–7.

Lu, C., Qing, Z., Li-Qiang, Z., and Yun, C. (2015). A trust-level based authorization model using signaling games. In 2015 IEEE 12th Intl Conf on Ubiquitous Intelligence and Computing and 2015 IEEE 12th Intl Conf on Autonomic and Trusted Computing and 2015 IEEE 15th Intl Conf on Scalable Computing and Communications and Its Associated Workshops (UIC-ATC-ScalCom), pages 772–778.

OWASP Foundation (2025). LLM01:2025 Prompt Injection. OWASP GenAI Security Project. Acesso em: 13 ago. 2026.

Pappu, K., Bhushan, B., and Mittal, A. (2025). Spiffe-based zero-trust authentication for ai agent ecosystems. In 2025 International Conference on Computer and Applications (ICCA), pages 1–7.

Pawlick, J., Colbert, E., and Zhu, Q. (2019). Modeling and analysis of leaky deception using signaling games with evidence. IEEE Transactions on Information Forensics and Security, 14(7):1871–1886.

Singh, A., Ehtesham, A., Lambe, M., Grogan, J., Singh, A., Kumar, S., Muscariello, L., Pandey, V., De Saint Marc, G. S., Chari, P., and Raskar, R. (2025). Evolution of ai agent registry solutions: Centralized, enterprise, and distributed approaches. In 2025 IEEE 7th International Conference on Cognitive Machine Intelligence (CogMI), pages 507–515.

Stefanescu, S. and Aciobănit,ei, I. (2026). Lineage: A uma2.0 protocol extension for multi-hop llm agent delegation. IEEE Access, 14:104574–104591.

Vassilev, A., Oprea, A., Fordyce, A., Anderson, H., Davies, X., and Hamin, M. (2025). Adversarial machine learning: A taxonomy and terminology of attacks and mitigations. Technical Report NIST AI 100-2e2025, National Institute of Standards and Technology.
Publicado
01/09/2026
BENTO, Lucila M. S.; BOCCARDO, Davidson R.. Jogos de Sinalização para Agentes de IA Comprometidos. In: WORKSHOP DE CIBERSEGURANÇA EM IA - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 1048-1056. DOI: https://doi.org/10.5753/sbseg_estendido.2026.33825.