Avaliação de Desempenho na Inferência de LLMs: Um Estudo Comparativo entre TPU e GPU utilizando vLLM
Resumo
Este estudo apresenta uma análise de desempenho focada na inferência de Grandes Modelos de Linguagem (LLMs), comparando a arquitetura Tensor Processing Unit (TPU) com aceleradores GPU tradicionais. A avaliação utiliza o modelo Llama-3.1 e o framework de inferência vLLM. Os resultados demonstram que a TPU supera a GPU de referência em throughput total, além de apresentar tempos de resposta (TTFT e TPOT) aproximadamente 2,5 vezes mais rápidos.Referências
Agrawal, A. et al. (2023). Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills. [link].
Jouppi, N. P. et al. (2017). In-datacenter performance analysis of a tensor processing unit. In Proc. of 44th ISCA, pages 1–12.
Kwon, W. et al. (2023). Efficient memory management for large language model serving with pagedattention. In Proc. of the 29th SOSP, page 611–626.
Nikolić, G. S. et al. (2022). A survey of three types of processing units: Cpu, gpu and tpu. In Proc. of the 57th ICEST, pages 1–6.
Silva, G. P., Bianchini, C. P., and Costa, E. B. (2022). Programação Paralela e Distribuída com MPI, OpenMP e OpenACC para computação de alto desempenho. CasaDoCodigo.
Touvron, H. et al. (2023). Llama: Open and efficient foundation language models. [link].
Jouppi, N. P. et al. (2017). In-datacenter performance analysis of a tensor processing unit. In Proc. of 44th ISCA, pages 1–12.
Kwon, W. et al. (2023). Efficient memory management for large language model serving with pagedattention. In Proc. of the 29th SOSP, page 611–626.
Nikolić, G. S. et al. (2022). A survey of three types of processing units: Cpu, gpu and tpu. In Proc. of the 57th ICEST, pages 1–6.
Silva, G. P., Bianchini, C. P., and Costa, E. B. (2022). Programação Paralela e Distribuída com MPI, OpenMP e OpenACC para computação de alto desempenho. CasaDoCodigo.
Touvron, H. et al. (2023). Llama: Open and efficient foundation language models. [link].
Publicado
02/09/2026
Como Citar
LIMA, Jose Augusto M. de; BIANCHINI, Calebe P..
Avaliação de Desempenho na Inferência de LLMs: Um Estudo Comparativo entre TPU e GPU utilizando vLLM. In: ESCOLA REGIONAL DE ALTO DESEMPENHO DE SÃO PAULO (ERAD-SP), 17. , 2026, São Paulo/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 33-36.
DOI: https://doi.org/10.5753/eradsp.2026.30932.
