Classification of Descriptive Sufficiency in Public Expenditure Liquidation Records

  • Jonathas E. Fujii UFMT
  • Gabriel H. T. Moreira UFMT
  • Thiago M. Ventura UFMT
  • Patícia C. de Souza UFMT
  • Raphael S. R. Gomes UFMT

Resumo


This study proposes a supervised Natural Language Processing approach based on BERTimbau to classify public expenditure liquidation records according to the descriptive sufficiency of their content. The research is grounded on the premise that generic, vague, or overly document-dependent descriptions reduce transparency and hinder institutional and social oversight of public spending. A labeled dataset of 600 records was constructed by random sampling with artificial class balancing from a corpus of 4,083 permanent equipment liquidation records from the State Executive Branch of Mato Grosso, Brazil, for fiscal year 2025, with labeling validated by two independent auditors via Cohen’s Kappa (κ = 0.900). Three supervised strategies were evaluated under stratified 10-fold cross-validation: TF-IDF + SVM (baseline), BERTimbau Feature-Based, and BERTimbau Fine-Tuning. The results demonstrate that both BERTimbau configurations outperformed the baseline with medium-to-large effect sizes, significant for Feature-Based, with Fine-Tuning achieving macro-F1 of 92.66% and accuracy of 92.67%, confirming the feasibility of contextualized Portuguese-language models for supporting preventive, risk-focused governmental auditing.

Palavras-chave: government auditing, bertimbau, text classification, public expenditure, natural language processing

Referências

Brasil. Lei n.º 4.320, de 17 de março de 1964. estatui normas gerais de direito financeiro para elaboração e controle dos orçamentos e balanços da união, dos estados, dos municípios e do distrito federal, 1964.

Dietterich, T. G. Approximate statistical tests for comparing supervised classification learning algorithms. Neural Computation 10 (7): 1895–1923, 1998.

Fantini, W. d. S. Publicação de Dados Conectados sobre Despesas Orçamentárias do Governo Federal Brasileiro. M.S. thesis, Centro de Informática, Universidade Federal de Pernambuco, Recife, PE, Brasil, 2015.

Kotepuchai, P. and Limpiyakorn, Y. Tree-based classifiers for smart general ledger code suggestion. In Proceedings of the 2024 13th International Conference on Informatics, Environment, Energy and Applications (IEEA 2024). ACM, Tokyo, Japan, pp. 17–22, 2024.

Kotios, D., Makridis, G., Fatouros, G., and Kyriazis, D. Deep learning enhancing banking services: a hybrid transaction classification and cash flow prediction approach. Journal of Big Data 9 (1): 130, Oct., 2022.

Michener, G., Contreras, E., and Niskier, I. Da opacidade à transparência? avaliando a Lei de Acesso à Informação no Brasil cinco anos depois. Revista de Administração Pública, 2018. RAP, Rio de Janeiro. Recebido set. 2017; aceito abr. 2018. [versão traduzida].

Santana, I. N., Oliveira, R. S., and Nascimento, E. G. S. Text classification of news using transformer-based models for portuguese. Journal of Systemics, Cybernetics and Informatics 20 (5): 33–59, 2022.

Santos, L. F. d. Aplicação de Machine Learning na Classificação de Restos a Pagar. M.S. thesis, Escola de Administração de Empresas de São Paulo, Fundação Getulio Vargas, São Paulo, SP, Brasil, 2016.

Silva, L. A. C., Rodrigues, M. S. F., Archanjo, A. P., Pessoa, L., Silva, M. L., Almeida, T. F., and Silveira, L. Segmentação textual baseada em tópicos em português utilizando bertimbau. In Anais do XV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL). SBC, Porto Alegre, RS, Brasil, pp. 32–36, 2024a.

Silva, M. O., Oliveira, G. P., Costa, L. G. L., and Pappa, G. L. Evaluating domain-adapted language models for governmental text classification tasks in Portuguese. In Anais do XXXIX Simpósio Brasileiro de Bancos de Dados (SBBD 2024). Sociedade Brasileira de Computação, Florianópolis, SC, Brasil, pp. 247–259, 2024b.

Souza, F. C., Nogueira, R. F., and Lotufo, R. A. BERT models for Brazilian Portuguese: Pretraining, evaluation and tokenization analysis. Applied Soft Computing vol. 149, pp. 110901, 2023.

Teixeira, P. H., da Silva, N. F. F., and Salvini, R. Machine learning based method for auditing personnel expenses in public expenditure. In Proceedings of the 2023 IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, Torino, Italy, pp. 320–327, 2023.

Vale, A. H., Santos, P., Soares, H., and Moura, R. S. Automatic classification of public expenses in the fight against COVID-19: A case study of TCE/PI. In Proceedings of the XIX Brazilian Symposium on Information Systems (SBSI 2023). ACM, Maceió, AL, Brasil, pp. 118–129, 2023.
Publicado
19/10/2026
FUJII, Jonathas E.; MOREIRA, Gabriel H. T.; VENTURA, Thiago M.; SOUZA, Patícia C. de; GOMES, Raphael S. R.. Classification of Descriptive Sufficiency in Public Expenditure Liquidation Records. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 81-88. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.31129.