Evaluating SLMs for Predicting Tabular Data: An Essay on Higher Education Dropout
Resumo
Student dropout in higher education represents a structural challenge for educational institutions, resulting in both academic and financial losses. Although classical machine learning models are successfully applied to this task, the consolidation of Large Language Models (LLMs) has enabled new predictive approaches. This study investigates the effectiveness of small-scale LLMs (7-14B parameters) adapted via Parameter-Efficient Fine-Tuning (PEFT) compared to specialized tabular-data models. Utilizing a UFPel dataset with 6,011 records, we evaluated three representation strategies: scaled (traditional), labeled (discretized), and storytelling (narrative), inferred by the Phi-4, Qwen2.5-7B, and Mitra models. The results indicate that the specialized tabular model (Mitra) outperformed the LLMs, achieving 80.30% accuracy and an F1-Score of 0.8203, compared to 72.07% and 0.7176 for the best-performing LLM (Phi-4 scaled). The labeled and storytelling linguistic formats yielded performance levels comparable to simple numerical formatting, invalidating the narrative advantage hypothesis in static prediction. Evidence shows that 4-bit-quantized small-scale LLMs perform worse on isolated tabular tasks. We conclude that tree-based and foundationally tabular classifiers remain the most effective architecture for dropout prediction under resource-constrained conditions.
Palavras-chave:
Small Language Models, Tabular Data, Higher Education Dropout
Referências
Armengol-Estapé, J., Woodruff, J., Cummins, C., and O'Boyle, M. F. P. (2024). SLaDe: A portable small language model decompiler for optimized assembly. In 2024 IEEE/ACM International Symposium on Code Generation and Optimization (CGO), pages 67–80.
Borisov, V., Seiffert, T., Mayrhauser, H., Molnar, C., and Kasneci, G. (2022). Deep learning for tabular data: An extensive survey. IEEE Transactions on Neural Networks and Learning Systems.
Daniel Han, M. H. and team, U. (2023). Unsloth.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized llms.
Do, T., Shrestha Gurung, B. D., Aryal, S., Khanal, A., Chataut, S., Gadhamshetty, V., Lushbough, C., and Gnimpieba, E. Z. (2023). Utilizing XGBoost for the Prediction of Material Corrosion Rates from Embedded Tabular Data using Large Language Model. In 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 4497–4499.
Fang, X., Xu, W., Tan, F. A., Zhang, J., Hu, Z., Qi, Y., Nickleach, S., Socolinsky, D., Sengamedu, S., and Faloutsos, C. (2024). Large language models (llms) on tabular data: Prediction, generation, and understanding - a survey. Transactions on Machine Learning Research.
Fukuda, N., Nozue, H., and Oishi, H. (2025). Small language model agent for the operations of continuously updating ICT systems. IEEE access : practical innovations, open solutions, pages 37522–37533.
Han, S., Yoon, J., Arik, S. O., and Pfister, T. (2024). Large language models can automatically engineer features for few-shot tabular learning. In Proceedings of the 41st International Conference on Machine Learning (ICML), pages 1–26.
Hosseini, E., Srinivas, A., Nazari, N., Hale, C., Rafatirad, S., and Homayoun, H. (2025). Large language models for opioid-induced respiratory depression prediction in hospitalized patients: A retrospective study. ACM Transactions on Computing for Healthcare, 6(2).
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR).
Hu, F., Fang, Y., Si, W., Dou, J., Yang, Y., Qiang, Y., and Dong, X. (2025). SLM-MEEF: ECG report generation based on small language model and multi-expert ensemble framework. In Proceedings of the 44th Chinese Control Conference (CCC), pages 8743–8748.
Kesarkar, A. G., Deepakbhai Patel, T., and Vinayak Deshmukh, P. (2025). Hybrid Intelligence in Business Analytics: Integrating Large Language Models with Traditional Predictive Analytics. In 2025 3rd International Conference on Sustainable Computing and Data Communication Systems (ICSCDS), pages 114–122.
Koska, B. and Horváth, M. (2024). Towards multi-modal mastery: A 4.5B parameter truly multi-modal small language model. In 2024 2nd International Conference on Foundation and Large Language Models (FLLM), pages 587–592.
Li, M., Xu, X., Li, S., and Wu, B. (2025). Making Small Language Model Excellent Symptom Inference Expert for Mental Disorders Detection. In 2025 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6, Nantes, France. IEEE.
Liang, V. W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J. (2022). Mind the gap: Understanding the modality gap in multi-modal contrastive learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, pages 17612–17625.
Ma, X., Liu, W., Zhao, C., and Tukhvatulina, L. R. (2024). Can Large Language Model Predict Employee Atrition? In Proceeding of the 2024 5th International Conference on Computer Science and Management Technology, pages 1164–1172, Xiamen Guangdong China. ACM.
Paul, S., Zhang, L., Shen, Y., and Jin, H. (2024). Enabling device control planning capabilities of small language model. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 12066–12070. IEEE.
Prasad, R., Singh, P. P., B, L., Maurya, A., D, S. N., and Singh, A. (2025). Sink vulnerability type prediction using small language model (SLM). In 2025 International Conference on Cognitive Computing in Engineering, Communications, Sciences and Biomedical Health Informatics (IC3ECSBHI), pages 496–501. IEEE.
Rabbani, S. B., Kowsar, I., and Samad, M. D. (2024). Transfer Learning of Tabular Data by Finetuning Large Language Models. In 2024 13th International Conference on Electrical and Computer Engineering (ICECE), pages 287–292.
Ridwansyah, A. and Saputro, I. P. (2025). Implementation of small language model for on-device creative lyrics generation using LoRA and 8-bit quantization methods. In 2025 International Conference on Informatics, Multimedia, Cyber and Information System (ICIMCIS), pages 552–555.
Thakur, S. K. and Tyagi, N. (2024). Spatial data discovery using small language model. In 2024 International Conference on Computing, Power, and Communication Technologies (IC2PCT), pages 899–905.
Zhang, X. and Robinson, D. M. (2025). Mitra: Mixed synthetic priors for enhancing tabular foundation models.
Borisov, V., Seiffert, T., Mayrhauser, H., Molnar, C., and Kasneci, G. (2022). Deep learning for tabular data: An extensive survey. IEEE Transactions on Neural Networks and Learning Systems.
Daniel Han, M. H. and team, U. (2023). Unsloth.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized llms.
Do, T., Shrestha Gurung, B. D., Aryal, S., Khanal, A., Chataut, S., Gadhamshetty, V., Lushbough, C., and Gnimpieba, E. Z. (2023). Utilizing XGBoost for the Prediction of Material Corrosion Rates from Embedded Tabular Data using Large Language Model. In 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 4497–4499.
Fang, X., Xu, W., Tan, F. A., Zhang, J., Hu, Z., Qi, Y., Nickleach, S., Socolinsky, D., Sengamedu, S., and Faloutsos, C. (2024). Large language models (llms) on tabular data: Prediction, generation, and understanding - a survey. Transactions on Machine Learning Research.
Fukuda, N., Nozue, H., and Oishi, H. (2025). Small language model agent for the operations of continuously updating ICT systems. IEEE access : practical innovations, open solutions, pages 37522–37533.
Han, S., Yoon, J., Arik, S. O., and Pfister, T. (2024). Large language models can automatically engineer features for few-shot tabular learning. In Proceedings of the 41st International Conference on Machine Learning (ICML), pages 1–26.
Hosseini, E., Srinivas, A., Nazari, N., Hale, C., Rafatirad, S., and Homayoun, H. (2025). Large language models for opioid-induced respiratory depression prediction in hospitalized patients: A retrospective study. ACM Transactions on Computing for Healthcare, 6(2).
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR).
Hu, F., Fang, Y., Si, W., Dou, J., Yang, Y., Qiang, Y., and Dong, X. (2025). SLM-MEEF: ECG report generation based on small language model and multi-expert ensemble framework. In Proceedings of the 44th Chinese Control Conference (CCC), pages 8743–8748.
Kesarkar, A. G., Deepakbhai Patel, T., and Vinayak Deshmukh, P. (2025). Hybrid Intelligence in Business Analytics: Integrating Large Language Models with Traditional Predictive Analytics. In 2025 3rd International Conference on Sustainable Computing and Data Communication Systems (ICSCDS), pages 114–122.
Koska, B. and Horváth, M. (2024). Towards multi-modal mastery: A 4.5B parameter truly multi-modal small language model. In 2024 2nd International Conference on Foundation and Large Language Models (FLLM), pages 587–592.
Li, M., Xu, X., Li, S., and Wu, B. (2025). Making Small Language Model Excellent Symptom Inference Expert for Mental Disorders Detection. In 2025 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6, Nantes, France. IEEE.
Liang, V. W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J. (2022). Mind the gap: Understanding the modality gap in multi-modal contrastive learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, pages 17612–17625.
Ma, X., Liu, W., Zhao, C., and Tukhvatulina, L. R. (2024). Can Large Language Model Predict Employee Atrition? In Proceeding of the 2024 5th International Conference on Computer Science and Management Technology, pages 1164–1172, Xiamen Guangdong China. ACM.
Paul, S., Zhang, L., Shen, Y., and Jin, H. (2024). Enabling device control planning capabilities of small language model. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 12066–12070. IEEE.
Prasad, R., Singh, P. P., B, L., Maurya, A., D, S. N., and Singh, A. (2025). Sink vulnerability type prediction using small language model (SLM). In 2025 International Conference on Cognitive Computing in Engineering, Communications, Sciences and Biomedical Health Informatics (IC3ECSBHI), pages 496–501. IEEE.
Rabbani, S. B., Kowsar, I., and Samad, M. D. (2024). Transfer Learning of Tabular Data by Finetuning Large Language Models. In 2024 13th International Conference on Electrical and Computer Engineering (ICECE), pages 287–292.
Ridwansyah, A. and Saputro, I. P. (2025). Implementation of small language model for on-device creative lyrics generation using LoRA and 8-bit quantization methods. In 2025 International Conference on Informatics, Multimedia, Cyber and Information System (ICIMCIS), pages 552–555.
Thakur, S. K. and Tyagi, N. (2024). Spatial data discovery using small language model. In 2024 International Conference on Computing, Power, and Communication Technologies (IC2PCT), pages 899–905.
Zhang, X. and Robinson, D. M. (2025). Mitra: Mixed synthetic priors for enhancing tabular foundation models.
Publicado
05/10/2026
Como Citar
SCAGLIONI, Fabrício G.; AGUIAR, Marilton; MATTOS, Júlio C. B..
Evaluating SLMs for Predicting Tabular Data: An Essay on Higher Education Dropout. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 987-1000.
DOI: https://doi.org/10.5753/sbie.2026.27434.
