Avaliação Automatizada de Feedbacks Gerados por LLMs em Questões de Biologia: Uma Abordagem via LLM-as-Judge

  • Ana Raissa Nascimento Silva Cesar School
  • Anderson Pinheiro Cavalcanti Cesar School / Universidade Federal Rural de Pernambuco (UFRPE)
  • Rafael Ferreira Mello Cesar School / Universidade Federal Rural de Pernambuco (UFRPE) https://orcid.org/0000-0003-3548-9670

Resumo


Oferecer feedback formativo em larga escala é um desafio, sobretudo em disciplinas factualmente densas como a Biologia. Este trabalho aplica uma abordagem LLM-as-judge a feedbacks gerados por três Large Language Models (LLMs), GPT-5.2, Gemini-2.5-Pro e Llama3.2:3B, sob três estratégias de prompt, em 50 respostas curtas de Biologia. Dois LLMs classificaram os 450 feedbacks em cinco dimensões de conteúdo. Os modelos diferiram significativamente, mas com efeitos pequenos: o GPT-5.2 mostrou-se analítico-crítico; o Gemini-2.5-Pro, motivacional; o Llama3.2:3B, fraco em alinhamento com objetivos. Identificou-se ainda viés de autoavaliação no GPT-5.2 e concordância apenas moderada entre avaliadores, evidenciando limites da avaliação automatizada por um único LLM.
Palavras-chave: LLM-as-Judge, Feedback Educacional, Avaliação Automatizada

Referências

Bewersdorff, A. et al. (2025). Towards adaptive feedback with AI: Comparing the feedback quality of LLMs and teachers on experimentation protocols. In arXiv preprint arXiv:2502.12842.

Cavalcanti, A. P., Barbosa, A., Carvalho, R., Freitas, F., Tsai, Y.-S., Gasevic, D., and Mello, R. F. (2021a). Automatic feedback in online learning environments: A systematic literature review. Computers and Education: Artificial Intelligence, 2:100027.

Cavalcanti, A. P., Mello, R. F., Miranda, P., Nascimento, A., and Freitas, F. (2021b). Utilização de recursos linguísticos para classificação automática de mensagens de feedback. In Anais do XXXII Simposio Brasileiro de Informática na Educação, pages 861–872. SBC.

Cohn, C. et al. (2024). Chain-of-thought prompting for automating feedback generation in programming education. arXiv preprint arXiv:2401.06554.

Dai, W., Tsai, Y.-S., Lin, J., Aldino, A., Jin, H., Li, T., Gasevic, D., and Chen, G. (2024). Assessing the proficiency of large language models in automatic feedback generation: An evaluation study. Computers and Education: Artificial Intelligence, 7:100299.

Galhardi, L. B., Senefonte, H., and Brancher, J. D. (2020). Portuguese automatic short answer grading. In Anais do Simposio Brasileiro de Informática na Educação (SBIE), pages 1373–1382.

Hattie, J. and Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1):81–112.

Lee, P. et al. (2025). Automated assignment grading with large language models: insights from a bioinformatics course. BMC Bioinformatics. Available at [link].

Lobo, J., Anthony, L., Falcao, A., Xavier, C., Torrezão, N., Isotani, S., Bittencourt, I. I., Rodrigues, L., and Mello, R. F. (2025). Automatic scoring of elementary school essays in brazilian portuguese with LLMs: Comparing gemini, GPT4o, claude, and mistral. In Anais do XXXVI Simposio Brasileiro de Informática na Educação (SBIE 2025), pages 168–181.

Mello, R. F., Freitas, E., Cabral, L., Pereira, F. D., Rodrigues, L., Rakovic, M., Raniel, J., and Gasevic, D. (2024). Words of wisdom: A journey through the realm of natural language processing for learning analyticsa systematic literature review. Journal of Learning Analytics, 11(3):82–105.

Nicol, D. J. and Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: A model and seven principles of good feedback practice. Studies in Higher Education, 31(2):199–218.

Oketch, K., Lalor, J. P., Yang, Y., and Abbasi, A. (2025). Bridging the LLM accessibility divide? performance, fairness, and cost of closed versus open LLMs for automated essay scoring. In Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEMˆ2). arXiv:2503.11827.

Pereira, A. F. and Ferreira Mello, R. (2025). A systematic literature review on large language models applications in computer programming teaching evaluation process. IEEE Access, 13:113449–113460.

Pimentel, E. P., da Silva, L. D. L., Takeuti, R. H., and Braga, J. C. (2025). Abordagem com LLM para gerac ̧ao automatizada de feedback no ensino de programação de computadores. In Anais do Simposio Brasileiro de Informática na Educação (SBIE), pages 590–602.

Qian, K., Cheng, Y., Guan, R., Dai, W., Jin, F., Yang, K., Nawaz, S., Swiecki, Z., Chen, G., Yan, L., and Gasevic, D. (2025). Dean of LLM tutors: Exploring comprehensive and automated evaluation of LLM-generated educational feedback via LLM feedback evaluators. arXiv preprint arXiv:2508.05952.

Silva, F. G. and da Silva Aranha, E. H. (2025). Feedback formativo automatizado com llms: Desenvolvimento e analise de um sistema para aprendizagem progressiva em programação. In Simposio Brasileiro de Informática na Educação (SBIE), pages 1158–1172. SBC.

Xavier, C., da Costa, N. T., Valdo, A. K., Alves, G., Rodrigues, L., Rodrigues, L. F., Silva, M., Neto, R., Falcao, T. P., Gasevic, D., et al. (2025). Human teacher vs. LLM-generated feedback in secondary education: A comparative study on student perceptions. In European Conference on Technology Enhanced Learning, pages 534–548.

Yan, L. et al. (2024). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1):90–112.
Publicado
05/10/2026
SILVA, Ana Raissa Nascimento; CAVALCANTI, Anderson Pinheiro; MELLO, Rafael Ferreira. Avaliação Automatizada de Feedbacks Gerados por LLMs em Questões de Biologia: Uma Abordagem via LLM-as-Judge. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 1732-1745. DOI: https://doi.org/10.5753/sbie.2026.28069.