ENEM Essay Feedback Using LLM-Augmented Prompts
Resumo
This paper evaluates automated feedback generation for ENEM Competency 5 by augmenting zero-shot prompts with an Evaluator Term Set (CTA), a term set proposed in this work and extracted from human evaluator comments on Competency 5. We compare CTA-augmented feedback against a zero-shot baseline on 46 essays with human reference feedback using BERTScore F1 and the Wilcoxon signed-rank test. CTA augmentation increases mean BERTScore F1 for both models, with statistically significant improvement in semantic similarity to human feedback for both evaluated models at α = 0.05.Referências
Anchiêta, R. T., Luz, A. I., Lopes, S. L., and Moura, R. S. (2025). A zero-shot prompting approach for automated feedback generation on ENEM essays. In Brazilian Symposium on Multimedia and the Web (WebMedia), pages 511–515. SBC.
Bezerra, E. (2025). Introduction to LLM-Based agents. In Duarte, D., Timbó, F., Schreiner, G., and Chaves, I., editors, Tópicos em Gerenciamento de Dados e Informações: Minicursos do SBBD 2025, chapter 2. Sociedade Brasileira de Computação, Porto Alegre.
Bird, S., Klein, E., and Loper, E. (2009). Natural language processing with Python: analyzing text with the natural language toolkit. O’Reilly Media, Inc.
Hooshyar, D., Yang, Y., Šíř, G., Kärkkäinen, T., Hämäläinen, R., Cukurova, M., and Azevedo, R. (2025). Problems with large language models for learner modelling: Why LLMs alone fall short for responsible tutoring in K–12 education. arXiv preprint arXiv:2512.23036.
INEP, I. (2024). ENEM 2024 tem 4,3 milhões de inscritos confirmados. Acesso em: 2 nov. 2024.
INEP, I. (2025). A redação no ENEM 2025: cartilha do(a) participante. Publicado em: 03 out. 2025 08:46. Acesso em: 30 mar. 2026.
Marcondes, F. S., Gala, A., Magalhães, R., Perez de Britto, F., Durães, D., and Novais, P. (2025). Using ollama. In Natural Language Analytics with Generative Large-Language Models: A Practical Approach with Ollama and Open-Source LLMs, pages 23–35. Springer.
MCTI, M. (2024). Plano brasileiro de inteligência artificial (PBIA) 2024-2028. Gov.br. Acesso em: 26 set. 2024.
Medeiros, M. (2016). Income inequality in Brazil: new evidence from combined tax and survey data. World Social Science Report, page 107.
Misgna, H., On, B.-W., Lee, I., and Choi, G. S. (2024). A survey on deep learning-based automated essay scoring and feedback generation. Artificial Intelligence Review, 58(2):36.
Perneger, T. V. (1998). What’s wrong with bonferroni adjustments. Bmj, 316(7139):1236–1238.
Schreiter, D. (2025). Prompt engineering: How prompt vocabulary affects domain knowledge. arXiv preprint arXiv:2505.17037.
Senkevics, A. S. and Carvalho, M. P. D. (2023). Youth and access to higher education: On the non-place of the preparatory course student. Educação Em Revista, 39.
Silva, W. A. d. and Araujo, C. C. d. (2024). Automated ENEM essay scoring and feedbacks: A prompt-driven LLM approach. Trabalho de Graduação, Centro de Informática (CIn), Universidade Federal de Pernambuco (UFPE), Recife, Brasil.
Silveira, I. C., Barbosa, A., and Mauá, D. D. (2024). A new benchmark for automatic essay scoring in Portuguese. In Proceedings of the 16th International Conference on Computational Processing of Portuguese-Vol. 1, pages 228–237.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2019). Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
Bezerra, E. (2025). Introduction to LLM-Based agents. In Duarte, D., Timbó, F., Schreiner, G., and Chaves, I., editors, Tópicos em Gerenciamento de Dados e Informações: Minicursos do SBBD 2025, chapter 2. Sociedade Brasileira de Computação, Porto Alegre.
Bird, S., Klein, E., and Loper, E. (2009). Natural language processing with Python: analyzing text with the natural language toolkit. O’Reilly Media, Inc.
Hooshyar, D., Yang, Y., Šíř, G., Kärkkäinen, T., Hämäläinen, R., Cukurova, M., and Azevedo, R. (2025). Problems with large language models for learner modelling: Why LLMs alone fall short for responsible tutoring in K–12 education. arXiv preprint arXiv:2512.23036.
INEP, I. (2024). ENEM 2024 tem 4,3 milhões de inscritos confirmados. Acesso em: 2 nov. 2024.
INEP, I. (2025). A redação no ENEM 2025: cartilha do(a) participante. Publicado em: 03 out. 2025 08:46. Acesso em: 30 mar. 2026.
Marcondes, F. S., Gala, A., Magalhães, R., Perez de Britto, F., Durães, D., and Novais, P. (2025). Using ollama. In Natural Language Analytics with Generative Large-Language Models: A Practical Approach with Ollama and Open-Source LLMs, pages 23–35. Springer.
MCTI, M. (2024). Plano brasileiro de inteligência artificial (PBIA) 2024-2028. Gov.br. Acesso em: 26 set. 2024.
Medeiros, M. (2016). Income inequality in Brazil: new evidence from combined tax and survey data. World Social Science Report, page 107.
Misgna, H., On, B.-W., Lee, I., and Choi, G. S. (2024). A survey on deep learning-based automated essay scoring and feedback generation. Artificial Intelligence Review, 58(2):36.
Perneger, T. V. (1998). What’s wrong with bonferroni adjustments. Bmj, 316(7139):1236–1238.
Schreiter, D. (2025). Prompt engineering: How prompt vocabulary affects domain knowledge. arXiv preprint arXiv:2505.17037.
Senkevics, A. S. and Carvalho, M. P. D. (2023). Youth and access to higher education: On the non-place of the preparatory course student. Educação Em Revista, 39.
Silva, W. A. d. and Araujo, C. C. d. (2024). Automated ENEM essay scoring and feedbacks: A prompt-driven LLM approach. Trabalho de Graduação, Centro de Informática (CIn), Universidade Federal de Pernambuco (UFPE), Recife, Brasil.
Silveira, I. C., Barbosa, A., and Mauá, D. D. (2024). A new benchmark for automatic essay scoring in Portuguese. In Proceedings of the 16th International Conference on Computational Processing of Portuguese-Vol. 1, pages 228–237.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2019). Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
Publicado
19/10/2026
Como Citar
CARVALHO, Flavio; SOARES, Vanessa; BEZERRA, Eduardo; GUEDES, Gustavo.
ENEM Essay Feedback Using LLM-Augmented Prompts. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 78-85.
DOI: https://doi.org/10.5753/stil.2026.26575.
