Juiz de Código Adaptativo: Uma Abordagem para Avaliação Justa de Algoritmos entre C++ e Python no Contexto Educacional
Resumo
Juízes de código online, amplamente usados no ensino de programação, adotam em geral tempos-limite fixos calibrados em C++. Esse modelo penaliza soluções Python de complexidade adequada com vereditos de Time Limit Exceeded (TLE) que refletem o overhead de execução, não a qualidade algorítmica. Este trabalho tem como objetivo promover equidade entre linguagens, propondo e validando um Juiz de Código Adaptativo, que calibra os tempos-limite por linguagem com benchmarks em contêineres Docker sobre classes sintéticas e onze problemas do CSES, com vereditos do juiz oficial. O limite adaptativo elimina o TLE injusto sem admitir soluções incorretas nem ineficientes, com fatores de 3× a 120× conforme o problema.
Referências
Beyer, D., Löwe, S., and Wendler, P. (2019). Reliable benchmarking: requirements and solutions. International Journal on Software Tools for Technology Transfer, 21(1):1–29.
Combéfis, S. (2022). Automated code assessment for education: Review, classification and perspectives on techniques and tools. Software, 1(1):3–30.
Efron, B. (1979). Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7(1):1–26.
Ihantola, P., Ahoniemi, T., Karavirta, V., and Seppälä, O. (2010). Review of recent systems for automatic assessment of programming assignments. In Proceedings of the 10th Koli Calling International Conference on Computing Education Research, pages 86–93.
Jain, R. (1991). The Art of Computer Systems Performance Analysis. John Wiley & Sons.
Kalibera, T. and Jones, R. (2013). Rigorous benchmarking in reasonable time. In Proceedings of the 2013 International Symposium on Memory Management, pages 63–74.
Laaksonen, A. (2017). Guide to Competitive Programming: Learning and Improving Algorithms Through Contests. Undergraduate Topics in Computer Science. Springer.
McGill, R., Tukey, J. W., and Larsen, W. A. (1978). Variations of box plots. The American Statistician, 32(1):12–16.
Merkel, D. (2014). Docker: lightweight linux containers for consistent development and deployment. Linux Journal, 2014(239):2.
Messer, M., Brown, N. C. C., Kölling, M., and Shi, M. (2024). Automated grading and feedback tools for programming education: A systematic review. ACM Transactions on Computing Education, 24(1):1–43.
Paiva, J. C., Leal, J. P., and Figueira, Á. (2022). Automated assessment in computer science education: A state-of-the-art review. ACM Transactions on Computing Education, 22(3):1–40.
Siegfried, R. M., Herbert-Berger, K. G., Leune, K., and Siegfried, J. P. (2021). Trends of commonly used programming languages in CS1 and CS2 learning. In Proceedings of the 16th International Conference on Computer Science & Education (ICCSE), pages 407–412. IEEE.
Silva, L. L. M. d., Cedraz, V. F., Patrocínio, J. A. d., Delabrida, S., and Fortes, R. S. (2024). Juiz online focado no ensino de programação: uma análise de usabilidade de ferramenta autoral. In Anais do XXXV Simposio Brasileiro de Informática na Educação (SBIE), pages 1234–1247, Porto Alegre. Sociedade Brasileira de Computação.
Wang, J., Lin, P., Tang, Z., and Chen, S. (2023). How problem difficulty and order influence programming education outcomes in online judge systems. Heliyon, 9(11).
Wasik, S., Antczak, M., Badura, J., Laskowski, A., and Sternal, T. (2018). A survey on online judge systems and their applications. ACM Computing Surveys, 51(1):1–34.
Zhang, J., Gu, J., Zhang, W., Cambronero, J. P., Kolesar, J., Piskac, R., Li, D., and Shi, H. (2025). A systematic study of time limit exceeded errors in online programming assignments. arXiv preprint arXiv:2510.14339.
