ITSM-BERT: Classificação de Tickets de Service Desk via Fine-Tuning do BERTimbau Integrado com Sistema GLPI
Resumo
Este artigo propõe o ITSM-BERT, um modelo baseado no BERTimbau, com ajuste fino para a classificação de chamados de Service Desk em português brasileiro, utilizando dados enriquecidos. A abordagem foi avaliada com 80.879 registros reais do Tribunal de Justiça do Maranhão, comparando métodos de frequência, vetoriais e Transformers (SVM-BoW, TF-IDF, Word2Vec e BERT-SVM) por meio de testes de Friedman e de Nemenyi. Na classificação autônoma, o ITSM-BERT supera todos os baselines (F1-macro = 0,840) e atinge 0,997 com atributos semiestruturados do catálogo de serviços (+15,7 p.p.). Embora estatisticamente equivalente ao SVM-TF-IDF no cenário assistido, consolida-se como uma solução robusta para mitigar ambiguidades semânticas em descrições brutas, com potencial para ser utilizada em outras empresas e órgãos públicos que utilizam sistemas de chamados, como o GLPI.
Palavras-chave:
BERTimbau, Classificação de Chamados, GLPI, ITSM, Service Desk
Referências
Alsaç, A., Yılmaz, Ü., Koçoğlu, F. Ö., and Şeker, Ş. E. (2025). Towards intelligent IT service management: A comparative evaluation of machine learning and language models for ticket classification. In UBMK, pages 265–270. IEEE.
Ayyalasomayajula, M. M. T., Bussa, S., Kowsalya, S. S. N., Mishra, N., Prasad, S., and Mehra, A. (2024). Leveraging feature extraction and fine-tuning techniques for enhanced detection of fake news using BERT. In ICAC2N, pages 1630–1635. IEEE.
Axelos (2019). ITIL Foundation. The Stationery Office, 4th edition.
Barchilon, N., Lopes, H. C. V., Kalinowski, M., and Perez, J. S. (2024). Enriquecimento de dados com base em estatísticas de grafo de similaridade para melhorar o desempenho em modelos de ML supervisionados de classificação. In Anais do XXXIX Simpósio Brasileiro de Banco de Dados (SBBD), pages 220–233. SBC.
Boonklay, J. and Jinarat, S. (2025). Comparative study of machine learning and deep learning algorithms for customer support ticket classification. In InCIT, pages 813–819. IEEE.
Bouckaert, R. R. and Frank, E. (2004). Evaluating the replicability of significance tests for comparing learning algorithms. In PAKDD, pages 3–12. Springer.
Brandão, M. A., Silva, M. O., Oliveira, G. P., Hott, H. R., Lacerda, A. M., and Pappa, G. L. (2023). Impacto do pré-processamento e representação textual na classificação de documentos de licitações. In Anais do XXXVIII Simpósio Brasileiro de Banco de Dados (SBBD), pages 102–114. SBC.
Carmo, F., Serejo, F., Jacob Junior, A., Santana, E., and Lobato, F. (2023). Embeddings jurídico: Representações orientadas à linguagem jurídica brasileira. In Anais do XI Workshop de Computação Aplicada em Governo Eletrônico, pages 188–199. SBC.
Chi, W. W., Tang, T. Y., Salleh, N. M., Mukred, M., AlSalman, H., and Zohaib, M. (2024). Data augmentation with semantic enrichment for deep learning invoice text classification. IEEE Access, 12:57326–57344.
Cui, Y., Jia, M., Lin, T.-Y., Song, Y., and Belongie, S. (2019). Class-balanced loss based on effective number of samples. In CVPR, pages 9268–9277. IEEE.
Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7(1):1–30.
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, pages 4171–4186.
Dlodlo, D. and Sibanda, K. (2023). Automated ticket classification for information technology helpdesks using machine learning. In ZCICT, pages 1–7. IEEE.
Ferdinand, Y., Lubis, M., and Pratiwi, O. N. (2025). A systematic literature review on AI and NLP applications for customer support automation and digital service. International Journal of Computer Technology and Science, 2(4):01–14.
Fuchs, S., Schnellbach, J., Wittges, H., and Krcmar, H. (2026). Human vs. automated data annotation: Labeling the data set for an ML-driven support ticket classifier. Data & Knowledge Engineering, DOI: 10.1016/j.datak.2025.102534.
ISACA (2019). COBIT 2019 Framework: Governance and Management Objectives. ISACA.
Jain, S., Gupta, A., and Neha, K. (2025). AI enhanced ticket management system for optimized support. In AIMLSystems, pages 1–7. ACM.
Kapoor, S. and Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9):100804.
Liu, F., He, X., Zhang, T., Chen, J., Li, Y., Yi, L., Zhang, H., Wu, G., and Shi, R. (2025). TickIt: Leveraging large language models for automated ticket escalation. In FSE, pages 343–354. ACM.
Liu, Z., Benge, C., and Jiang, S. (2023). Ticket-BERT: Labeling incident management tickets with language models. arXiv preprint arXiv:2307.00108.
Paket, E., Şenerkek, G., Akyol, F. B., and Salman, F. (2024). IT service desk ticket classification via large language models. In UBMK, pages 135–140. IEEE.
Paramesh, S. P. and Shreedhara, K. S. (2019). Automated IT service desk systems using machine learning techniques. In Data Analytics and Learning, Lecture Notes in Networks and Systems, volume 43, pages 331–346. Springer.
Sant'Anna, Y. F. D., Souza, L. F. C., Alves Neto, A. J., Rodrigues Junior, M. C., Carvalho, A. B., and Gusmão, R. P. (2024). Classificação da dívida ativa do estado de Sergipe. In Anais do XXXIX Simpósio Brasileiro de Banco de Dados (SBBD), pages 562–573. SBC.
Silva, M. O., Oliveira, G. P., Costa, L. G. L., and Pappa, G. L. (2024). Evaluating domain-adapted language models for governmental text classification tasks in Portuguese. In Anais do XXXIX Simpósio Brasileiro de Banco de Dados (SBBD), pages 247–259. SBC.
Souza, F., Nogueira, R., and Lotufo, R. (2020). BERTimbau: Pretrained BERT models for Brazilian Portuguese. In BRACIS, pages 403–417. Springer.
Souza, F. and Souza Filho, J. B. O. (2023). Embedding generation for text classification of Brazilian Portuguese user reviews: From bag-of-words to transformers. Neural Computing and Applications, 35(13):9393–9406.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. In NeurIPS, volume 30, pages 5998–6008.
Wirth, R. and Hipp, J. (2000). CRISP-DM: Towards a standard process model for data mining. In Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining, volume 1, pages 29–39.
Ayyalasomayajula, M. M. T., Bussa, S., Kowsalya, S. S. N., Mishra, N., Prasad, S., and Mehra, A. (2024). Leveraging feature extraction and fine-tuning techniques for enhanced detection of fake news using BERT. In ICAC2N, pages 1630–1635. IEEE.
Axelos (2019). ITIL Foundation. The Stationery Office, 4th edition.
Barchilon, N., Lopes, H. C. V., Kalinowski, M., and Perez, J. S. (2024). Enriquecimento de dados com base em estatísticas de grafo de similaridade para melhorar o desempenho em modelos de ML supervisionados de classificação. In Anais do XXXIX Simpósio Brasileiro de Banco de Dados (SBBD), pages 220–233. SBC.
Boonklay, J. and Jinarat, S. (2025). Comparative study of machine learning and deep learning algorithms for customer support ticket classification. In InCIT, pages 813–819. IEEE.
Bouckaert, R. R. and Frank, E. (2004). Evaluating the replicability of significance tests for comparing learning algorithms. In PAKDD, pages 3–12. Springer.
Brandão, M. A., Silva, M. O., Oliveira, G. P., Hott, H. R., Lacerda, A. M., and Pappa, G. L. (2023). Impacto do pré-processamento e representação textual na classificação de documentos de licitações. In Anais do XXXVIII Simpósio Brasileiro de Banco de Dados (SBBD), pages 102–114. SBC.
Carmo, F., Serejo, F., Jacob Junior, A., Santana, E., and Lobato, F. (2023). Embeddings jurídico: Representações orientadas à linguagem jurídica brasileira. In Anais do XI Workshop de Computação Aplicada em Governo Eletrônico, pages 188–199. SBC.
Chi, W. W., Tang, T. Y., Salleh, N. M., Mukred, M., AlSalman, H., and Zohaib, M. (2024). Data augmentation with semantic enrichment for deep learning invoice text classification. IEEE Access, 12:57326–57344.
Cui, Y., Jia, M., Lin, T.-Y., Song, Y., and Belongie, S. (2019). Class-balanced loss based on effective number of samples. In CVPR, pages 9268–9277. IEEE.
Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7(1):1–30.
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, pages 4171–4186.
Dlodlo, D. and Sibanda, K. (2023). Automated ticket classification for information technology helpdesks using machine learning. In ZCICT, pages 1–7. IEEE.
Ferdinand, Y., Lubis, M., and Pratiwi, O. N. (2025). A systematic literature review on AI and NLP applications for customer support automation and digital service. International Journal of Computer Technology and Science, 2(4):01–14.
Fuchs, S., Schnellbach, J., Wittges, H., and Krcmar, H. (2026). Human vs. automated data annotation: Labeling the data set for an ML-driven support ticket classifier. Data & Knowledge Engineering, DOI: 10.1016/j.datak.2025.102534.
ISACA (2019). COBIT 2019 Framework: Governance and Management Objectives. ISACA.
Jain, S., Gupta, A., and Neha, K. (2025). AI enhanced ticket management system for optimized support. In AIMLSystems, pages 1–7. ACM.
Kapoor, S. and Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9):100804.
Liu, F., He, X., Zhang, T., Chen, J., Li, Y., Yi, L., Zhang, H., Wu, G., and Shi, R. (2025). TickIt: Leveraging large language models for automated ticket escalation. In FSE, pages 343–354. ACM.
Liu, Z., Benge, C., and Jiang, S. (2023). Ticket-BERT: Labeling incident management tickets with language models. arXiv preprint arXiv:2307.00108.
Paket, E., Şenerkek, G., Akyol, F. B., and Salman, F. (2024). IT service desk ticket classification via large language models. In UBMK, pages 135–140. IEEE.
Paramesh, S. P. and Shreedhara, K. S. (2019). Automated IT service desk systems using machine learning techniques. In Data Analytics and Learning, Lecture Notes in Networks and Systems, volume 43, pages 331–346. Springer.
Sant'Anna, Y. F. D., Souza, L. F. C., Alves Neto, A. J., Rodrigues Junior, M. C., Carvalho, A. B., and Gusmão, R. P. (2024). Classificação da dívida ativa do estado de Sergipe. In Anais do XXXIX Simpósio Brasileiro de Banco de Dados (SBBD), pages 562–573. SBC.
Silva, M. O., Oliveira, G. P., Costa, L. G. L., and Pappa, G. L. (2024). Evaluating domain-adapted language models for governmental text classification tasks in Portuguese. In Anais do XXXIX Simpósio Brasileiro de Banco de Dados (SBBD), pages 247–259. SBC.
Souza, F., Nogueira, R., and Lotufo, R. (2020). BERTimbau: Pretrained BERT models for Brazilian Portuguese. In BRACIS, pages 403–417. Springer.
Souza, F. and Souza Filho, J. B. O. (2023). Embedding generation for text classification of Brazilian Portuguese user reviews: From bag-of-words to transformers. Neural Computing and Applications, 35(13):9393–9406.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. In NeurIPS, volume 30, pages 5998–6008.
Wirth, R. and Hipp, J. (2000). CRISP-DM: Towards a standard process model for data mining. In Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining, volume 1, pages 29–39.
Publicado
08/09/2026
Como Citar
MAIA DE LIMA CARVALHO, Anderson; FRANÇA LOBATO, Fábio Manoel; LAVAREDA JACOB JUNIOR, Antonio Fernando; CARMONA CORTES, Omar Andres.
ITSM-BERT: Classificação de Tickets de Service Desk via Fine-Tuning do BERTimbau Integrado com Sistema GLPI. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 317-330.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249215.
