Geração de casos de teste com IA Generativa para plataformas de juiz online: um estudo comparativo com o Pynguin
Resumo
Juízes online são amplamente utilizados no ensino de programação, mas dependem de conjuntos de questões e testes abrangentes para avaliar soluções corretamente. Este trabalho compara a geração automática de testes com GPT-5.4 e Pynguin em 120 problemas do LeetCode, utilizando cobertura de código e mutation score como métricas. Com uma estratégia de self-refining, a abordagem baseada em IA alcançou 95% de cobertura em média e superou o Pynguin em 84% dos problemas em termos do mutation score. Os resultados sugerem que modelos de inteligência artificial generativa constituem uma alternativa promissora para a geração automática de testes, e que supera outro método de geração automática de testes.
Palavras-chave:
Geração de Casos de Teste, IA Generativa, Juiz Online
Referências
Bhatia, S., Gandhi, T., Kumar, D., and Jalote, P. (2024). Unit test generation using generative ai: A comparative performance analysis of autogeneration tools. In Proceedings of the 1st International Workshop on Large Language Models for Code, LLM4Code '24, page 54–61, New York, NY, USA. Association for Computing Machinery.
Cao, Y., Chen, Z., Quan, K., Zhang, Z., Wang, Y., Dong, X., Feng, Y., He, G., Huang, J., Li, J., Tan, Y., Tang, J., Tang, Y., Wu, J., Xiao, Q., Zheng, C., Zhou, S., Zhu, Y., Huang, Y., and He, T. (2026). Can llms generate reliable test case generators? a study on competition-level programming problems.
Carvalho, L., Oliveira, D., and Gadelha, B. (2016). Juiz online como ferramenta de apoio a uma metodologia de ensino híbrido em programação. In Anais do XXVII Simpósio Brasileiro de Informática na Educação, pages 140–149, Porto Alegre, RS, Brasil. SBC.
Chen, P.-Y. (2026). Leetcode solutions. [link]. Repositório GitHub. Acesso em: 31 maio 2026.
Dantas, J. and Aranha, E. (2025). Sistema multiagente para a geração automática de questões de programação. In Anais Estendidos do XIV Congresso Brasileiro de Informática na Educação, pages 346–352, Porto Alegre, RS, Brasil. SBC.
Francisco, R., Júnior, C. P., and Ambrósio, A. P. (2016). Juiz online no ensino de programação introdutória - uma revisão sistemática da literatura. In Anais do XXVII Simpósio Brasileiro de Informática na Educação, pages 11–20, Porto Alegre, RS, Brasil. SBC.
Guilherme, V. and Vincenzi, A. (2023). An initial investigation of chatgpt unit test generation capability. In Anais do VIII Simpósio Brasileiro de Testes de Software Sistemático e Automatizado, page 15–24, Porto Alegre, RS, Brasil. SBC.
Lukasczyk, S. and Fraser, G. (2022). Pynguin: automated unit test generation for python. In Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings, ICSE '22, page 168–172, New York, NY, USA. Association for Computing Machinery.
Oshin, M. and Campos, N. (2025). Learning LangChain: Building AI and LLM Applications with LangChain and LangGraph. O'Reilly Media, Inc., Sebastopol, CA, USA. Copyright © 2025 Olumayowa "Mayo" Olufemi Oshin. All rights reserved. Printed in the United States of America.
Santana, A., Silva, F. G., Dantas, J., Souza, J., and Aranha, E. (2025a). Geração automática de questões de programação usando LLM: Um relato de experiência. In Anais do XXXIII Workshop sobre Educação em Computação (WEI), pages 1415–1425, Maceió, AL, Brasil. Sociedade Brasileira de Computação.
Santana, A., Silva, F. G., Souza, J. L. G., da Silva Dantas, J. C., and Aranha, E. (2025b). Geração de questões de programação baseada em templates e ia generativa. In Workshop de Informática na Escola, CBIE.
Shah, C. (2024). From prompt engineering to prompt science with human in the loop. In Proceedings of ACM Conference on Human in the Loop, Seattle, WA, USA. arXiv preprint arXiv:2401.04122, May 10, 2024.
Silva, E., Coelho, R., and Silva, L. (2025). Llms as test generators: A comparative benchmarking study. In Anais do XXXIX Simpósio Brasileiro de Engenharia de Software, pages 25–36, Porto Alegre, RS, Brasil. SBC.
Tang, Y., Liu, Z., Zhou, Z., and Luo, X. (2024). Chatgpt vs sbst: A comparative assessment of unit test suite generation. IEEE Transactions on Software Engineering.
Yuan, Z., Liu, M., Ding, S., Wang, K., Chen, Y., Peng, X., and Lou, Y. (2024). Evaluating and improving chatgpt for unit test generation. Proceedings of the ACM on Software Engineering, 1(FSE).
Cao, Y., Chen, Z., Quan, K., Zhang, Z., Wang, Y., Dong, X., Feng, Y., He, G., Huang, J., Li, J., Tan, Y., Tang, J., Tang, Y., Wu, J., Xiao, Q., Zheng, C., Zhou, S., Zhu, Y., Huang, Y., and He, T. (2026). Can llms generate reliable test case generators? a study on competition-level programming problems.
Carvalho, L., Oliveira, D., and Gadelha, B. (2016). Juiz online como ferramenta de apoio a uma metodologia de ensino híbrido em programação. In Anais do XXVII Simpósio Brasileiro de Informática na Educação, pages 140–149, Porto Alegre, RS, Brasil. SBC.
Chen, P.-Y. (2026). Leetcode solutions. [link]. Repositório GitHub. Acesso em: 31 maio 2026.
Dantas, J. and Aranha, E. (2025). Sistema multiagente para a geração automática de questões de programação. In Anais Estendidos do XIV Congresso Brasileiro de Informática na Educação, pages 346–352, Porto Alegre, RS, Brasil. SBC.
Francisco, R., Júnior, C. P., and Ambrósio, A. P. (2016). Juiz online no ensino de programação introdutória - uma revisão sistemática da literatura. In Anais do XXVII Simpósio Brasileiro de Informática na Educação, pages 11–20, Porto Alegre, RS, Brasil. SBC.
Guilherme, V. and Vincenzi, A. (2023). An initial investigation of chatgpt unit test generation capability. In Anais do VIII Simpósio Brasileiro de Testes de Software Sistemático e Automatizado, page 15–24, Porto Alegre, RS, Brasil. SBC.
Lukasczyk, S. and Fraser, G. (2022). Pynguin: automated unit test generation for python. In Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings, ICSE '22, page 168–172, New York, NY, USA. Association for Computing Machinery.
Oshin, M. and Campos, N. (2025). Learning LangChain: Building AI and LLM Applications with LangChain and LangGraph. O'Reilly Media, Inc., Sebastopol, CA, USA. Copyright © 2025 Olumayowa "Mayo" Olufemi Oshin. All rights reserved. Printed in the United States of America.
Santana, A., Silva, F. G., Dantas, J., Souza, J., and Aranha, E. (2025a). Geração automática de questões de programação usando LLM: Um relato de experiência. In Anais do XXXIII Workshop sobre Educação em Computação (WEI), pages 1415–1425, Maceió, AL, Brasil. Sociedade Brasileira de Computação.
Santana, A., Silva, F. G., Souza, J. L. G., da Silva Dantas, J. C., and Aranha, E. (2025b). Geração de questões de programação baseada em templates e ia generativa. In Workshop de Informática na Escola, CBIE.
Shah, C. (2024). From prompt engineering to prompt science with human in the loop. In Proceedings of ACM Conference on Human in the Loop, Seattle, WA, USA. arXiv preprint arXiv:2401.04122, May 10, 2024.
Silva, E., Coelho, R., and Silva, L. (2025). Llms as test generators: A comparative benchmarking study. In Anais do XXXIX Simpósio Brasileiro de Engenharia de Software, pages 25–36, Porto Alegre, RS, Brasil. SBC.
Tang, Y., Liu, Z., Zhou, Z., and Luo, X. (2024). Chatgpt vs sbst: A comparative assessment of unit test suite generation. IEEE Transactions on Software Engineering.
Yuan, Z., Liu, M., Ding, S., Wang, K., Chen, Y., Peng, X., and Lou, Y. (2024). Evaluating and improving chatgpt for unit test generation. Proceedings of the ACM on Software Engineering, 1(FSE).
Publicado
05/10/2026
Como Citar
DANTAS, Júlio César da S.; SILVA, Francisco Genivan; SOUZA, Jadson Lucas Gomes; ARANHA, Eduardo Henrique da Silva.
Geração de casos de teste com IA Generativa para plataformas de juiz online: um estudo comparativo com o Pynguin. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 2191-2204.
DOI: https://doi.org/10.5753/sbie.2026.28258.
