Beyond Efficiency: The Impact of Instance Selection on Semantic Locality in Transformer-based Text Classification
Resumo
Transformer-based models achieve high performance in text classification tasks. In parallel, Instance Selection (IS) techniques have been widely adopted to reduce training sets by removing redundant or noisy examples, lowering computational cost while preserving predictive effectiveness. However, little is known about how these techniques affect the semantic organization of the representation space learned by Transformer models. This work investigates the impact of two recent IS methods, E2SC and biO-IS, on the semantic locality structure of BERT representations and on locality properties previously associated with the quality of model explanations. Models were trained on the original AGNews dataset and on reduced training sets produced by each IS method. Semantic locality was characterized using CLS embeddings from the last BERT layer, average neighborhood similarity and class entropy computed over k-Nearest Neighbors (kNN), complemented by a qualitative analysis based on neighborhood word clouds. The results reveal a dissociated effect of Instance Selection on locality: while both methods increase neighborhood similarity, only biO-IS consistently increases the concentration of instances in highly homogeneous regions. These findings indicate that different Instance Selection strategies modify the geometry of the learned representation space in distinct ways, suggesting that their effects extend beyond computational efficiency and may influence locality properties relevant to explanation methods.
Referências
Cunha, W., Canuto, S., Viegas, F., Salles, T., Gomes, C., Mangaravite, V., Resende, E., Rosa, T., Gonçalves, M. A., and Rocha, L. Extended pre-processing pipeline for text classification: On the role of meta-feature representations, sparsification and selective sampling. Information Processing & Management 57 (4): 102263, 2020.
Cunha, W., Mangaravite, V., Gomes, C., Canuto, S., Resende, E., Viegas, F., Martins, W. S., Almeida, J. M., et al. On the cost-effectiveness of neural and non-neural approaches and representations for text classification: A comprehensive comparative study. Information Processing & Management 58 (3): 102481, 2021.
Cunha, W., Rocha, L., and Gonçalves, M. A. A Noise-Oriented and Redundancy-Aware Instance Selection Framework. ACM Transactions on Information Systems 43 (2): 1–33, 2025.
Cunha, W., Viegas, F., França, C., Rosa, T., Rocha, L., and Gonçalves, M. A. A comparative survey of instance selection methods applied to non-neural and transformer-based text classification. ACM CSUR, 2023.
Cunha, W., Viegas, F., França, C., Rosa, T., Rocha, L., and Gonçalves, M. A. An Effective, Efficient, and Scalable Confidence-Based Instance Selection Framework for Transformer-Based Text Classification. In Proceedings of the 46th International ACM SIGIR. Taipei, Taiwan, pp. 665–674, 2023.
de Andrade, C. M., Belém, F., Cunha, W., França, C., Viegas, F., Rocha, L., and Gonçalves, M. A. On the class separability of contextual embeddings representations – or “the classifier does not matter when the (text) representation is so good!”. Information Processing & Management 60 (4): 103336, 2023.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the NAACL. pp. 4171–4186, 2019.
Fernandes, D., de Moura, E. S., Ribeiro-Neto, B. A., da Silva, A. S., and Gonçalves, M. A. Computing block importance for searching on web sites. In CIKM ’07. pp. 165–174, 2007.
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint 1 (1): 1–13, 2019.
Maharana, A., Yadav, P., and Bansal, M. D2 pruning: Message passing for balancing diversity and difficulty in data pruning. In International Conference on Learning Representations (ICLR), 2024.
Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. JMLR 12 (1): 2825–2830, 2011. Rahulamathavan, Y., Farooq, M., and Silva, V. D. PLEX: Perturbation-Free Local Explanations for LLM-Based Text Classification. arXiv preprint 1 (1): 1–15, 2025.
Ribeiro, M. T., Singh, S., and Guestrin, C. Why Should I Trust You? Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD. San Francisco, USA, pp. 1135–1144, 2016.
Sener, O. and Savarese, S. Active learning for convolutional neural networks: A core-set approach. In International Conference on Learning Representations (ICLR), 2018.
Soares, L. S. et al. Are all instances equally explainable? a study on the impact of locality on the explainability of automatic text classification tasks. In Proceedings of the World Conference on xAI, 2026.
Swayamdipta, S., Schwartz, R., Lourie, N., Wang, Y., Hajishirzi, H., Smith, N. A., and Choi, Y. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In EMNLP. pp. 9275–9293, 2020.
Viegas, F., Cunha, W., Gomes, C., Pereira, A., Rocha, L., and Goncalves, M. CluHTM - semantic hierarchical topic modeling based on CluWords. In Proceedings of the 58th ACL. Online, pp. 8138–8150, 2020.
Wolf, T. et al. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Online, pp. 38–45, 2020.
Zanotto, B. S., Beck da Silva Etges, A. P., Dal Bosco, A., Cortes, E. G., Ruschel, R., De Souza, A. C., Andrade, C. M., Viegas, F., Luiz, W., et al. Stroke outcome measurements from electronic medical records: cross-sectional study on the effectiveness of neural and nonneural classifiers. JMIR Med. Inf. 9 (11): e29120, 2021.
Zhang, X., Zhao, J., and LeCun, Y. Character-Level Convolutional Networks for Text Classification. In Advances in Neural Information Processing Systems. Montreal, Canada, pp. 649–657, 2015.
