The Role of Retained-Layer Representations in Data Selection for Compressed Transformer Fine-Tuning

  • Victoria F. Mello UFMG
  • Flavio Soriano UFMG
  • Pedro B. Rigueira UFMG
  • Gisele L. Pappa UFMG
  • Wagner Meira Jr. UFMG
  • Washington Cunha Unicamp
  • Marcos André Gonçalves UFMG

Resumo


Fine-tuning of Transformer-based models can be made more efficient by reducing the model, the training set, or both. However, layer pruning and training-instance selection are often treated as independent decisions, even though pruning layers changes the representation space in which examples are evaluated. This paper investigates whether data selection for fine-tuning compressed Transformer-based models should account for the representation space defined by the retained layers. We evaluate this hypothesis using BERT-base, selecting layers through kernel–target alignment (KTA) to favor task-relevant representations and centered kernel alignment (CKA) to penalize redundancy. The selected layers then define the representation kernel used for Nyström-based training-instance selection, followed by candidate-pool construction and balanced downselection. In experiments with BERT-base on SST-2 and QNLI across 12 seeds, the evaluated configuration achieves the highest average accuracy among the compressed variants considered in our experimental setting. On QNLI, it reaches (61.36±6.44)% accuracy and improves over BI+Nyström-Uniform by (6.25) paired percentage points, with 9/11 wins and BH-FDR (¯0.0205). On SST-2, it reaches (79.88 ± 6.08)% accuracy and remains competitive among the compressed configurations. Across both tasks, it reduces the aggregate FLOPs proxy by about 69% and peak CUDA memory by 13–19% relative to full BERT, although selection overhead still limits wall-clock gains. These results provide initial evidence that the representation space defined by the retained layers is a relevant consideration for data selection in compressed BERT fine-tuning.

Palavras-chave: centered kernel alignment, efficient fine-tuning, layer-aware data selection, Nyström approximation, pruning

Referências

Cortes, C., Mohri, M., and Rostamizadeh, A. Algorithms for learning kernels based on centered alignment. Journal of Machine Learning Research 13 (28): 795–828, 2012.

Cunha, W., França, C., Fonseca, G., Rocha, L., and Gonçalves, M. A. An effective, efficient, and scalable confidence-based instance selection framework for transformer-based text classification. In SIGIR. pp. 665–674, 2023.

Cunha, W., Moreo Fernández, A., Esuli, A., Sebastiani, F., Rocha, L., and Gonçalves, M. A. A noise-oriented and redundancy-aware instance selection framework. ACM TOIS 43 (2): 1–33, 2025.

Cunha, W., Rocha, L., and Gonçalves, M. A. Inteligência artificial sustentável baseado em engenharia de dados, aprendizado de máquina e transferência de conhecimento para processamento de linguagem natural. In SBBD, 2025.

Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the NAACL. pp. 4171–4186, 2019.

Di Liello, L., Gabburo, M., and Moschitti, A. Effective pretraining objectives for transformer-based autoencoders. In Findings of EMNLP 2022. Association for Computational Linguistics, pp. 5533–5547, 2022.

El Alaoui, A. and Mahoney, M. W. Fast randomized kernel ridge regression with statistical guarantees. In Advances in Neural Information Processing Systems. Vol. 28. pp. 775–783, 2015.

Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. Similarity of neural network representations revisited. In ICML. Vol. 97. PMLR, pp. 3519–3529, 2019.

Margatina, K., Barrault, L., and Aletras, N. On the importance of effectively adapting pretrained language models for active learning. In ACL. pp. 825–836, 2022.

Men, X., Xu, M., Zhang, Q., Wang, B., Lin, H., Lu, Y., Han, X., and Chen, W. ShortGPT: Layers in large language models are more redundant than you expect. In Findings of ACL 2025, 2025.

Mosbach, M., Andriushchenko, M., and Klakow, D. On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines, 2021.

Musco, C. and Musco, C. Recursive sampling for the Nyström method. In NeurIPS. Vol. 30, 2017.

Paul, M., Ganguli, S., and Dziugaite, G. K. Deep learning on a data diet: Finding important examples early in training. In Advances in Neural Information Processing Systems. Vol. 34. pp. 20596–20607, 2021.

Sanh, V., Wolf, T., and Rush, A. M. Movement pruning: Adaptive sparsity by fine-tuning. In Advances in Neural Information Processing Systems. Vol. 33. pp. 20378–20389, 2020.

Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP 2013. pp. 1631–1642, 2013.

Toneva, M., Sordoni, A., Tachet des Combes, R., Trischler, A., Bengio, Y., and Gordon, G. J. An empirical study of example forgetting during deep neural network learning. In ICLR, 2019.

Viegas, F., Cunha, W., Gomes, C., Pereira, A., Rocha, L., and Goncalves, M. CluHTM - semantic hierarchical topic modeling based on CluWords. In Proceedings of the 58th ACL. pp. 8138–8150, 2020.

Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In ICLR, 2019.

Williams, C. K. I. and Seeger, M. Using the Nyström method to speed up kernel machines. In Advances in Neural Information Processing Systems. Vol. 13. MIT Press, pp. 682–688, 2001.

Xia, M., Zhong, Z., and Chen, D. Structured pruning learns compact and accurate models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. pp. 1513–1528, 2022.

Xu, C., Zhou, W., Ge, T., Wei, F., and Zhou, M. BERT-of-Theseus: Compressing BERT by progressive module replacing. In Proceedings of the 2020 EMNLP. Association for Computational Linguistics, pp. 7859–7869, 2020.
Publicado
19/10/2026
MELLO, Victoria F.; SORIANO, Flavio; RIGUEIRA, Pedro B.; PAPPA, Gisele L.; MEIRA JR., Wagner; CUNHA, Washington; GONÇALVES, Marcos André. The Role of Retained-Layer Representations in Data Selection for Compressed Transformer Fine-Tuning. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 241-248. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.32102.