Avaliação de LLMs como Recurso de Apoio ao Aprendizado de Programação Básica

  • Mika M. dos Santos Júnior Universidade Federal do Maranhão (UFMA)
  • Alanna C. da Silva Centro Universitário Unidade de Ensino Superior Dom Bosco (UNDB)
  • Diego M. A. Oliveira Universidade Federal do Maranhão (UFMA)
  • Djefferson Maranhão Universidade Federal do Maranhão (UFMA)
  • Carlos de Salles Soares Neto Universidade Federal do Maranhão (UFMA)

Resumo


O crescente uso de grandes modelos de linguagem (LLMs) tem impactado plataformas colaborativas como o Stack Overflow, reduzindo o cadastro de novas perguntas ao longo dos anos, o que gera dúvidas sobre a qualidade dessas respostas. Diante disso, este estudo realiza uma avaliação da qualidade semântica, por meio da métrica BERTScore, da capacidade dos modelos GPT-4.1 Mini, Gemini 2.5 Flash, Claude Sonnet 4.6 e DeepSeek-V3 em responder perguntas de programação, usando os 90 pares de pergunta e resposta mais bem avaliados da plataforma como benchmark. Os resultados mostram similaridade semântica moderada, indicando que esses modelos têm potencial, mas ainda não substituem o suporte humano no aprendizado de programação.
Palavras-chave: Modelos de Linguagem de Grande Porte, BERTScore, Aprendizado de Programação

Referências

Anthropic (2026). Claude sonnet 4.6. [link]. Acesso em: 15 mar. 2026.

Banerjee, S. and Lavie, A. (2005). Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.

da Cunha, M. V., Silveira, M. R., Santana, B. S., Freitas, L. A., and Corrêa, U. B. (2025). Optimizing and evaluating a retrieval-augmented generation system for normative document retrieval in hospital settings. In Brazilian Symposium on Multimedia and the Web (WebMedia), pages 385–393. SBC.

da Silva Junior, S. M., de Freitas, R. A. B., de Morais, M. A. C., and Costa, D. L. V. (2023). Chatgpt no auxílio da aprendizagem de programação: Um estudo de caso. In Simpósio Brasileiro de Informática na Educação (SBIE), pages 1375–1384. SBC.

del Rio-Chanona, M., Laurentsyeva, N., and Wachs, J. (2023). Are large language models a threat to digital public goods? evidence from activity on stack overflow. arXiv preprint arXiv:2307.07367.

Denny, P., Prather, J., Becker, B. A., Finnie-Ansley, J., Hellas, A., Leinonen, J., Luxton-Reilly, A., Reeves, B. N., Santos, E. A., and Sarsa, S. (2024). Computing education in the era of generative ai. Communications of the ACM, 67(2):56–67.

Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171–4186.

Google Cloud (2025). Gemini 2.5 flash. [link]. Acessado em: 18 dez. 2025.

Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al. (2024). Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437.

Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.

Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., and Mian, A. (2024). A comprehensive overview of large language models. arXiv preprint arXiv:2307.06435.

OpenAI (2025). Gpt-4.1 mini. [link]. Acessado em: 27 mar. 2026.

Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.

Stack Exchange (2026). Stack Exchange API documentation. [link]. Acessado em: 23 dez. 2025.

Warner, B., Chaffin, A., Clavié, B., Weller, O., Hallström, O., Taghadouini, S., Gallagher, A., Biswas, R., Ladhak, F., Aarsen, T., et al. (2025). Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2526–2547.

Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2019). Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
Publicado
05/10/2026
DOS SANTOS JÚNIOR, Mika M.; DA SILVA, Alanna C.; OLIVEIRA, Diego M. A.; MARANHÃO, Djefferson; SOARES NETO, Carlos de Salles. Avaliação de LLMs como Recurso de Apoio ao Aprendizado de Programação Básica. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 2387-2400. DOI: https://doi.org/10.5753/sbie.2026.28463.