Caracterização da Inferência de LLMs Limitada por Memória em Hierarquias de DRAM e Memória Persistente

  • Luísa Leona Cattai UNESP
  • Alexandro Baldassin UNESP

Resumo


A inferência de Modelos de Linguagem de Grande Escala (LLMs) em ambientes exclusivamente baseados em CPU é fundamentalmente limitada pela capacidade e pela largura de banda da memória principal (DRAM). Este trabalho investiga a estratificação de memória (memory tiering) que distribui os três componentes internos do Transformer (pesos de atenção, pesos da rede Feed-Forward e KV-Cache) entre DRAM local e uma camada de maior capacidade e maior latência (Intel Optane DC). Sobre dois modelos LLaMA de 8B com arquiteturas de atenção distintas (MHA e GQA), constatamos que o componente dominante na degradação de vazão depende do regime de contexto: a FFN domina sob contextos curtos e o KV-Cache sob contextos longos. A partir desses achados, formulam-se cinco diretrizes práticas de alocação.

Referências

Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S. (2023). GQA: Training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the EMNLP, pages 4895–4901.

Alizadeh, K., Mirzadeh, I., Belenko, D., Khatamifard, S. K., Cho, M., Del Mundo, C. C., Rastegari, M., and Farajtabar, M. (2024). LLM in a flash: Efficient large language model inference with limited memory. In Proceedings of the 62nd ACL, pages 12562–12584.

Das Sharma, D., Blankenship, R., and Berger, D. (2024). An introduction to the compute express link (cxl) interconnect. ACM Comput. Surv., 56(11).

Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. (2023). OPTQ: Accurate quantization for generative pre-trained transformers. In Proceedings of the 11th ICLR.

Gerganov, G. et al. (2023). llama.cpp. [link]. Accessed: 2025.

Luo, C., Cai, Z., Sun, H., Xiao, J., Yuan, B., Xiao, W., Hu, J., Zhao, J., Chen, B., and Anandkumar, A. (2025). HeadInfer: Memory-efficient LLM inference by head-wise offloading. arXiv preprint arXiv:2502.12574.

Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Jin, H., Chen, T., and Jia, Z. (2024). Towards efficient generative large language model serving: A survey from algorithms to systems. arXiv preprint arXiv:2312.15234.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30.

Yang, J., Kim, J., Hoseinzadeh, M., Izraelevitz, J., and Swanson, S. (2020). An empirical guide to the behavior and use of scalable persistent memory. In 18th USENIX Conference on File and Storage Technologies (FAST ’20), pages 169–182.

Yuan, Z., Shang, Y., Zhou, Y., Dong, Z., Zhou, Z., Xue, C., Wu, B., Li, Z., Gu, Q., Lee, Y. J., Yan, Y., Chen, B., Sun, G., and Keutzer, K. (2024). LLM inference unveiled: Survey and roofline model insights. arXiv preprint arXiv:2402.16363.
Publicado
02/09/2026
CATTAI, Luísa Leona; BALDASSIN, Alexandro. Caracterização da Inferência de LLMs Limitada por Memória em Hierarquias de DRAM e Memória Persistente. In: ESCOLA REGIONAL DE ALTO DESEMPENHO DE SÃO PAULO (ERAD-SP), 17. , 2026, São Paulo/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 25-28. DOI: https://doi.org/10.5753/eradsp.2026.30867.

Artigos mais lidos do(s) mesmo(s) autor(es)

<< < 1 2 3