Análise de Desempenho e Quantização do ModernBERTBr
Resumo
Análise do impacto da redução de precisão numérica no modelo ModernBERTBr em tarefas de similaridade semântica (STS) e inferência lógica (RTE) em português. Utilizando uma GPU NVIDIA RTX 5070, foram testados formatos de alta precisão (FP32, FP16, BF16) e versões quantizadas (8 e 4 bits). Os resultados demonstram que o uso de BF16 e 4 bits reduz o consumo de VRAM em 50% e otimiza a latência sem comprometer a acurácia, que se manteve em 83,91% no RTE. Conclui-se que a quantização é uma técnica eficaz para viabilizar modelos de linguagem robustos em hardwares limitados.Referências
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. (2022). Gpt3. int8 (): 8bit matrix multiplication for transformers at scale. Advances in neural information processing systems, 35:30318–30332.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. Advances in neural information processing systems, 36:10088–10115.
Lang, J., Guo, Z., and Huang, S. (2024). A comprehensive study on quantization techniques for large language models. In 2024 4th International conference on artificial intelligence, robotics, and communication (ICAIRC), pages 224–231. IEEE.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
Warner, B., Chaffin, A., Clavié, B., Weller, O., Hallström, O., Taghadouini, S., Gallagher, A., Biswas, R., Ladhak, F., Aarsen, T., et al. (2025). Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2526–2547.
Wu, W. B. T. L. and Garcia, L. P. F. (2025). Modbertbr: A modernbert-based model for brazilian portuguese. In Encontro Nacional de Inteligência Artificial e Computacional (ENIAC), pages 2044–2055. SBC.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. Advances in neural information processing systems, 36:10088–10115.
Lang, J., Guo, Z., and Huang, S. (2024). A comprehensive study on quantization techniques for large language models. In 2024 4th International conference on artificial intelligence, robotics, and communication (ICAIRC), pages 224–231. IEEE.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
Warner, B., Chaffin, A., Clavié, B., Weller, O., Hallström, O., Taghadouini, S., Gallagher, A., Biswas, R., Ladhak, F., Aarsen, T., et al. (2025). Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2526–2547.
Wu, W. B. T. L. and Garcia, L. P. F. (2025). Modbertbr: A modernbert-based model for brazilian portuguese. In Encontro Nacional de Inteligência Artificial e Computacional (ENIAC), pages 2044–2055. SBC.
Publicado
02/09/2026
Como Citar
SOARES, Anderson Barbosa.
Análise de Desempenho e Quantização do ModernBERTBr. In: ESCOLA REGIONAL DE ALTO DESEMPENHO DE SÃO PAULO (ERAD-SP), 17. , 2026, São Paulo/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 93-96.
DOI: https://doi.org/10.5753/eradsp.2026.30995.
