Protegendo Modelos de IA com GPUs Confidenciais em Nuvens Não Confiáveis

  • Hipólito Araújo UFRN
  • Lucas Garcia UFRN
  • Leonardo Cavalcanti UFRN
  • Werbert Barradas UFRN
  • Eduardo Falcão UFRN
  • Andrey Brito UFCG

Resumo


À medida que modelos de Inteligência Artificial (IA) se tornam amplamente adotados, cresce a necessidade de mecanismos que garantam a confidencialidade e a integridade. Neste trabalho, projetamos e implementamos mecanismos que combinam o provisionamento de identidades e a computação confidencial para viabilizar IA confidencial em GPUs modernas com suporte à execução protegida. A solução estende a ferramenta de gerenciamento de identidades SPIRE com um atestador de nó capaz de verificar, em conjunto, CPU e GPU por meio do serviço de atestação remota da NVIDIA, emitindo identidades verificáveis (certificados X.509 ou JWTs) para cargas de trabalho do Kubernetes. Resultados experimentais em nuvens públicas demonstram que os mecanismos introduzem sobrecarga limitada, de aproximadamente 0,8 segundos, mantendo a integração transparente ao ecossistema nativo em nuvem.

Referências

Alibaba Cloud (2025). Build a heterogeneous confidential computing environment. [link].

AMD (2020). AMD SEV-SNP: Strengthening VM Isolation with Integrity Protection and More. Technical report.

Anthropic (2025). Confidential inference systems: Design principles and security risks. Technical report, Anthropic and Pattern Labs.

Birkholz, H., Thaler, D., Richardson, M., Smith, N., and Pan, W. (2023). Rfc 9334: Remote attestation procedures (rats) architecture.

Bodea, T., Misono, M., Pritzi, J., Sabanic, P., Sommer, T., Unnibhavi, H., Schall, D., Santos, N., Stavrakakis, D., and Bhatotia, P. (2025). Trusted ai agents in the cloud.

CoCo Contributors (2026a). Cloud api adaptor. [link]. Accessed: 2026-05-21.

CoCo Contributors (2026b). Confidential Containers. [link].

Consortium, C. C. et al. (2022). A technical analysis of confidential computing. Confidential Computing Consortium–Linux Foundation, Technical Report v1, 3.

Dhanuskodi, G., Guha, S., Krishnan, V., Manjunatha, A., O’Connor, M., Nertney, R., and Rogers, P. (2023). Creating the first confidential gpus: The team at nvidia brings confidentiality and integrity to user code and data for accelerated computing. Queue, 21(4):68–93.

Falcão, E., Silva, F., Pamplona, C., Melo, A., Asadujjaman, A. S. M., and Brito, A. (2025). Confidential kubernetes deployment models: Architecture, security, and performance trade-offs. Applied Sciences, 15(18).

Feldman, D., Fox, E., Gilman, E., Haken, I., Kautz, F., Khan, U., Lambrecht, M., Lum, B., Fayó, A. M., Nesterov, E., Vega, A., and Wardrop, M. (2020). Solving the Bottom Turtle: a SPIFFE way to establish trust in your infrastructure via universal identity.

Google (2026a). Confidential space overview. [link]. Accessed: 2026-05-28.

Google (2026b). Prompt encryption sdk. [link]. Accessed: 2026-05-28.

Google Cloud (2025). How confidential computing lays the foundation for trusted ai. [link].

Huang, C.-j., Zhao, H., He, Y., Li, L., Jiao, W., Jin, Z., Chen, P., and Wang, L. (2026). Your inference request will become a black box: Confidential inference for cloud-based large language models.

Intel (2025). Intel® Trust Domain Extensions (Intel® TDX) . Technical report.

Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I. (2023). Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pages 611–626.

Lan, H., Zhou, Z., Zhu, Q., Yan, W., Hao, Q., Ye, X., Liu, Y., and Sun, N. (2025). Heterogeneous confidential computing system for large language models: A survey. ACM Trans. Archit. Code Optim.

Li, C., Gim, I., and Zhong, L. (2024). Confidential prompting: Privacy-preserving llm inference on cloud.

Microsoft (2024). General availability: Azure confidential vms with nvidia h100 tensor core gpus. [link].

Narayan, A., Biderman, D., and Ré, C. (2025). Proof-of-concept for private local-to-cloud llm chat via trusted execution environments. In ICML 2025 Workshop on Efficient Systems for Foundation Models.

Neural Magic, I. (2024). Guidellm: Scalable inference and optimization for large language models. [link].

NVIDIA (2023a). Confidential computing: The developer’s view to secure an application and data on nvidia h100. [link].

NVIDIA (2023b). NVIDIA H100 GPU Whitepaper. [link].

Pontes, D., Silva, F., Falcão, E., and Brito, A. (2023). Attesting amd sev-snp virtual machines with spire. In Latin-American Symposium on Dependable and Secure Computing, page 1–10, New York, NY, USA. Association for Computing Machinery.

Pontes, D., Silva, F., Melo, A., Asadujjaman, A., Falcão, E., Brito, A., and Filho, C. (2024a). Multi-platform and vault-free attestation of confidential vms. In Latin-American Symposium on Dependable and Secure Computing, LADC ’24, page 241–251, New York, NY, USA. Association for Computing Machinery.

Pontes, D., Silva, F., Melo, A., Falcão, E., and Brito, A. (2024b). Interoperable node integrity verification for confidential machines based on amd sev-snp. Journal of Internet Services and Applications, 15(1):179–193.

SPIRE Contributors (2026). The spiffe runtime environment github project. [link].

Stanford Center for Research on Foundation Models (2026). Helm mmlu leaderboard. [link]. Accessed: 2026-05-21.

Sun, Y., Li, Y., Zhang, Y., Jin, Y., and Zhang, H. (2024). Svip: Towards verifiable inference of open-source large language models. arXiv preprint arXiv:2410.22307.

Vieira, M. (2025). Why we should trust systems, not just their ai/ml components. Computer, 58(11):84–94.
Publicado
01/09/2026
ARAÚJO, Hipólito; GARCIA, Lucas; CAVALCANTI, Leonardo; BARRADAS, Werbert; FALCÃO, Eduardo; BRITO, Andrey. Protegendo Modelos de IA com GPUs Confidenciais em Nuvens Não Confiáveis. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 977-992. DOI: https://doi.org/10.5753/sbseg.2026.27136.

Artigos mais lidos do(s) mesmo(s) autor(es)

1 2 > >>