Exploring Problem Statements and Source Code Submissions in Online Judges via 2D Embedding Projections
Resumo
Online Judges (OJs) are widely used as supporting tools in computer programming courses. These platforms store large collections of problem statements and student code submissions, offering rich opportunities for analyzing behaviors and performance patterns. However, their built-in analytics capabilities are limited, making systematic data exploration difficult. In this paper, we investigate how embeddings of problem statements and student source code, combined with multidimensional projection techniques, can support exploratory analysis of student activity. We evaluate multiple embedding models for Portuguese problem statements, as well as for source code, using metrics such as Neighborhood Preservation, Trustworthiness, Spearman correlation, and Stress. Our results indicate that ModernBERT-PT and GraphCodeBERT are the most suitable embeddings for Portuguese statements and source code, respectively. Regarding projection methods, t-SNE outperformed UMAP and TriMap in preserving local structure. Two case studies illustrate the practical value of the approach, showing how projections reveal groups of students with similar topic preferences and highlight potentially suspicious submissions.
Palavras-chave:
visual data mining, embeddings, multidimensional projection, online judges, educational data mining
Referências
Amid, E. and Warmuth, M. K. TriMap: Large-scale Dimensionality Reduction Using Triplets. arXiv preprint arXiv:1910.00204, 2019.
Carissimi, M., Saletta, M., and Ferretti, C. Towards leveraging large language model summaries for topic modeling in source code. In Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering. EASE ’25. Association for Computing Machinery, New York, NY, USA, pp. 776–781, 2025.
Feist, M. D., Santos, E. A., Watts, I., and Hindle, A. Visualizing project evolution through abstract syntax tree analysis. In 2016 IEEE Working Conference on Software Visualization (VISSOFT). IEEE, pp. 11–20, 2016.
Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., and Zhou, M. Codebert: A pre-trained model for programming and natural languages, 2020.
Günther, M., Ong, J., Mohr, I., Abdessalem, A., Abel, T., Akram, M. K., Guzman, S., Mastrapas, G., Sturua, S., Wang, B., et al. Jina embeddings 2: 8192-token general-purpose text embeddings for long documents. arXiv preprint arXiv:2310.19923 , 2023.
Guo, D., Ren, S., Lu, S., Feng, Z., Tang, D., Liu, S., Zhou, L., Duan, N., Svyatkovskiy, A., Fu, S., et al. Graphcodebert: Pre-training code representations with data flow. arXiv preprint arXiv:2009.08366 , 2020.
Keim, D., Andrienko, G., Fekete, J.-D., Görg, C., Kohlhammer, J., and Melançon, G. Visual analytics: Definition, process, and challenges. In Information Visualization: Human-centered Issues and Perspectives. Springer, pp. 154–175, 2008.
Kim, J., Cho, E., and Na, D. Problem-solving guide (psg): Predicting the algorithm tags and difficulty for competitive programming problems. In AI for Education Workshop. PMLR, pp. 48–56, 2024.
McInnes, L., Healy, J., and Melville, J. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 , 2018.
Mosquera, M. and Bojorque, R. Evaluating embedding representations for multiclass code smell detection: A comparative study of codebert and general-purpose embeddings. Applied Sciences 16 (8): 3622, 2026.
Nonato, L. G. and Aupetit, M. Multidimensional projection for visual analytics: Linking techniques with distortions, tasks, and layout enrichment. IEEE Transactions on Visualization and Computer Graphics 25 (8): 2650–2673, 2018.
Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). pp. 3982–3992, 2019.
Silvestre, A. S. S., De Souza, B. V., Lisboa, V. H. F., and Borges, V. R. P. A multi-label classification approach for categorizing beginner programming problems from online judges. In 2024 IEEE Frontiers in Education Conference (FIE). pp. 1–8, 2024.
Souza, F., Nogueira, R., and Lotufo, R. Bertimbau: Pretrained bert models for brazilian portuguese. In Intelligent Systems, R. Cerri and R. C. Prati (Eds.). Springer International Publishing, Cham, pp. 403–417, 2020.
Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of Machine Learning Research 9 (11): 2579–2605, 2008.
Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., and Wei, F. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533 , 2022.
Wang, Y., Le, H., Gotmare, A., Bui, N., Li, J., and Hoi, S. Codet5+: Open code large language models for code understanding and generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. pp. 1069–1088, 2023.
Wang, Z., Zhang, W., and Wang, J. Estimating difficulty levels of programming problems with pre-trained model. arXiv preprint arXiv:2406.08828 , 2024.
Yoshida, C., Matsushita, M., and Higo, Y. Estimating the Difficulty of Programming Problems Using Fine-tuned LLM . In IEEE/ACIS 22nd International Conference on Software Engineering Research, Management and Applications (SERA). pp. 28–34, 2024.
Carissimi, M., Saletta, M., and Ferretti, C. Towards leveraging large language model summaries for topic modeling in source code. In Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering. EASE ’25. Association for Computing Machinery, New York, NY, USA, pp. 776–781, 2025.
Feist, M. D., Santos, E. A., Watts, I., and Hindle, A. Visualizing project evolution through abstract syntax tree analysis. In 2016 IEEE Working Conference on Software Visualization (VISSOFT). IEEE, pp. 11–20, 2016.
Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., and Zhou, M. Codebert: A pre-trained model for programming and natural languages, 2020.
Günther, M., Ong, J., Mohr, I., Abdessalem, A., Abel, T., Akram, M. K., Guzman, S., Mastrapas, G., Sturua, S., Wang, B., et al. Jina embeddings 2: 8192-token general-purpose text embeddings for long documents. arXiv preprint arXiv:2310.19923 , 2023.
Guo, D., Ren, S., Lu, S., Feng, Z., Tang, D., Liu, S., Zhou, L., Duan, N., Svyatkovskiy, A., Fu, S., et al. Graphcodebert: Pre-training code representations with data flow. arXiv preprint arXiv:2009.08366 , 2020.
Keim, D., Andrienko, G., Fekete, J.-D., Görg, C., Kohlhammer, J., and Melançon, G. Visual analytics: Definition, process, and challenges. In Information Visualization: Human-centered Issues and Perspectives. Springer, pp. 154–175, 2008.
Kim, J., Cho, E., and Na, D. Problem-solving guide (psg): Predicting the algorithm tags and difficulty for competitive programming problems. In AI for Education Workshop. PMLR, pp. 48–56, 2024.
McInnes, L., Healy, J., and Melville, J. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 , 2018.
Mosquera, M. and Bojorque, R. Evaluating embedding representations for multiclass code smell detection: A comparative study of codebert and general-purpose embeddings. Applied Sciences 16 (8): 3622, 2026.
Nonato, L. G. and Aupetit, M. Multidimensional projection for visual analytics: Linking techniques with distortions, tasks, and layout enrichment. IEEE Transactions on Visualization and Computer Graphics 25 (8): 2650–2673, 2018.
Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). pp. 3982–3992, 2019.
Silvestre, A. S. S., De Souza, B. V., Lisboa, V. H. F., and Borges, V. R. P. A multi-label classification approach for categorizing beginner programming problems from online judges. In 2024 IEEE Frontiers in Education Conference (FIE). pp. 1–8, 2024.
Souza, F., Nogueira, R., and Lotufo, R. Bertimbau: Pretrained bert models for brazilian portuguese. In Intelligent Systems, R. Cerri and R. C. Prati (Eds.). Springer International Publishing, Cham, pp. 403–417, 2020.
Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of Machine Learning Research 9 (11): 2579–2605, 2008.
Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., and Wei, F. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533 , 2022.
Wang, Y., Le, H., Gotmare, A., Bui, N., Li, J., and Hoi, S. Codet5+: Open code large language models for code understanding and generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. pp. 1069–1088, 2023.
Wang, Z., Zhang, W., and Wang, J. Estimating difficulty levels of programming problems with pre-trained model. arXiv preprint arXiv:2406.08828 , 2024.
Yoshida, C., Matsushita, M., and Higo, Y. Estimating the Difficulty of Programming Problems Using Fine-tuned LLM . In IEEE/ACIS 22nd International Conference on Software Engineering Research, Management and Applications (SERA). pp. 28–34, 2024.
Publicado
19/10/2026
Como Citar
BORGES, Vinicius R. P.; HOLANDA, Maria Eduarda M. de; MOLCHANOV, Vladimir; OLIVEIRA, Maria Cristina F.; LINSEN, Lars.
Exploring Problem Statements and Source Code Submissions in Online Judges via 2D Embedding Projections. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 105-112.
ISSN 2763-8944.
DOI: https://doi.org/10.5753/kdmile.2026.32014.
