SLMs como Avaliadores Automatizados: Avaliando Competências de Escrita Narrativa com Modelos Adaptados por LoRA
Resumo
A avaliação de redações narrativas no ensino fundamental é pedagogicamente relevante, mas custosa, demorada e sujeita à variabilidade entre avaliadores. Como LLMs já foram estudados nesta tarefa, mas têm inferência custosa, este trabalho investiga se pequenos modelos de linguagem (SLMs), adaptados com LoRA e executados localmente com quantização de 4 bits, podem apoiar a correção automática de redações narrativas em português brasileiro. Foram avaliados seis modelos das famílias SmolLM, Gemma-3 e Qwen2.5, de 135M a 1B parâmetros, em um corpus anotado em quatro competências: registro formal, coerência temática, estrutura retórica narrativa e coesão. Comparamos modelos adaptados por LoRA com duas baselines zero-shot: o prompt original da tarefa e um prompt estrito orientado à geração de JSON válido. Os resultados mostram que o prompting zero-shot é pouco confiável para avaliação fim-a-fim: o prompt original não produziu saídas estruturadas válidas, enquanto o prompt estrito melhorou a validade em alguns modelos, mas manteve baixo desempenho de pontuação. Em contraste, todos os SLMs adaptados por LoRA geraram saídas válidas para todo o conjunto de teste. O melhor modelo, SmolLM-135M-4bit, obteve acurácia de 0,617, acerto off-by-one de 90,5% e MAE de 0,480. Os achados sugerem que SLMs quantizados e adaptados por LoRA são uma alternativa viável para avaliação automática de redações em contextos educacionais com restrição computacional.
Palavras-chave:
Avaliação automática de redações, Pequenos modelos de linguagem, LoRA
Referências
Allal, L. B., Lozhkov, A., Bakouch, E., Blázquez, G. M., Penedo, G., Tunstall, L., Marafioti, A., Kydlíček, H., Piqueres Lajarín, A., Srivastav, V., Lochner, J., Fahlgren, C., Nguyen, X.-S., Fourrier, C., Burtenshaw, B., Larcher, H., Zhao, H., Zakka, C., Morlon, M., Raffel, C., von Werra, L., and Wolf, T. (2025). Smollm2: When smol goes big – data-centric training of a fully open small language model. In Proceedings of the Conference on Language Modeling.
Barrera, S. D. and Santos, M. J. d. (2016). Produção escrita de narrativas: influência de condições de solicitação. Educar em Revista, (62):69–85.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. In Advances in Neural Information Processing Systems, volume 36, pages 10088–10115.
Ferreira, R., Freitas, E., Cabral, L., Dawn, F., Rodrigues, L., Rakovic, M., Raniel, J., and Gasevic, D. (2024). Words of wisdom: A journey through the realm of nlp for learning analytics – a systematic literature review. Journal of Learning Analytics, 11(3):82–105.
Gemma Team (2025). Gemma 3 technical report. arXiv preprint arXiv:2503.19786.
Graham, S. and Harris, K. R. (2019). Evidence-based practices in writing. In Graham, S., MacArthur, C. A., and Hebert, M., editors, Best Practices in Writing Instruction. The Guilford Press, 3 edition.
Hou, Z., Ciuba, A., and Li, X. (2025). Improving llm-based automatic essay scoring with linguistic features. In Proceedings of the Innovation and Responsibility in AI-Supported Education Workshop, volume 273 of Proceedings of Machine Learning Research, pages 41–65. PMLR.
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022). Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
Hugging Face TB (2024). Smollm-135m model card. [link]. Accessed: 2026-05-30.
Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., and Kasneci, G. (2023). Chatgpt for good? on opportunities and challenges of large language models for education. Learning and Individual Differences, 103:102274.
Liew, P. Y. and Tan, I. K. T. (2024). On automated essay grading using large language models. In Proceedings of the 2024 8th International Conference on Computer Science and Artificial Intelligence.
Liu, Z., Zhao, C., Iandola, F., Lai, C., Tian, Y., Fedorov, I., Xiong, Y., Chang, E., Shi, Y., Krishnamoorthi, R., Lai, L., and Chandra, V. (2024). Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 32431–32454. PMLR.
Lobo, J., Anthony, L., Falcão, A., Xavier, C., Torrezão, N., Isotani, S., Pinto, I. I. B. S., Rodrigues, L. A. L., and Mello, R. F. (2025). Automatic scoring of elementary school essays in brazilian portuguese with llms: Comparing gemini, gpt-4o, claude, and mistral. In Anais do Simpósio Brasileiro de Informática na Educação.
Lu, Z., Li, X., Cai, D., Yi, R., Liu, F., Zhang, X., and Lane, N. D. (2024). Small language models: Survey, measurements, and insights. arXiv preprint arXiv:2409.15790.
Mello, R. F., Oliveira, H., Wenceslau, M., Batista, H., Cordeiro, T., Bittencourt, I. I., and Isotani, S. (2024). Propor'24 competition on automatic essay scoring of portuguese narrative essays. In Proceedings of the 16th International Conference on Computational Processing of Portuguese, pages 1–5.
Nguyen, C. V., Shen, X., Aponte, R., Xia, Y., Basu, S., Hu, Z., Chen, J., Parmar, M., Kunapuli, S., Barrow, J., Wu, J., Singh, A., Wang, Y., Gu, J., Ahmed, N. K., Lipka, N., Zhang, R., Chen, X., Yu, T., Kim, S., Deilamsalehy, H., Park, N., Rimer, M., Zhang, Z., Yang, H., Mathur, P., Wu, G., Dernoncourt, F., Rossi, R. A., and Nguyen, T. H. (2025). A survey on small language models. In Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing, pages 807–821. INCOMA Ltd.
Oliveira, H., Mello, R. F., Miranda, P., Batista, H., Silva Filho, M. W. d., Cordeiro, T., Pinto, I. I. B. S., and Isotani, S. (2025). A benchmark dataset of narrative student essays with multi-competency grades for automatic essay scoring in brazilian portuguese. Data in Brief, 60:111526.
Page, E. B. (1966). The imminence of grading essays by computer. The Phi Delta Kappan, 47(5):238–243.
Qwen Team (2024). Qwen2.5 technical report. arXiv preprint arXiv:2412.15115.
Seßler, K., Fürstenberg, M., Bühler, B., and Kasneci, E. (2025). Can ai grade your essays? a comparative analysis of large language models and teacher ratings in multidimensional essay scoring. In Proceedings of the 15th International Conference on Learning Analytics and Knowledge, pages 462–472. Association for Computing Machinery.
Shermis, M. D. and Burstein, J. (2013). Handbook of Automated Essay Evaluation: Current Applications and New Directions. Routledge.
Siemens, G. (2013). Learning analytics: The emergence of a discipline. American Behavioral Scientist, 57(10):1380–1400.
Silva Filho, M. W. d. and collaborators (2024). Brazilian portuguese narrative essays dataset. [link]. Accessed: 2026-05-30.
Silva Filho, M. W. d., Nascimento, A., Miranda, P., Rodrigues, L. A. L., Cordeiro, T., Isotani, S., Pinto, I. I. B. S., and Mello, R. F. (2023). Automated formal register scoring of student narrative essays written in portuguese. In Anais do II Workshop de Aplicações Práticas de Learning Analytics em Instituições de Ensino no Brasil, pages 1–11. SBC.
Slade, S. and Prinsloo, P. (2013). Learning analytics: Ethical issues and dilemmas. American Behavioral Scientist, 57(10):1510–1529.
Wang, D. and Wang, J. (2025). The impact mechanism of aes on improving english writing achievement. Scientific Reports, 15:3928.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. (2026). A survey of large language models. Frontiers of Computer Science, 20:2012627.
Barrera, S. D. and Santos, M. J. d. (2016). Produção escrita de narrativas: influência de condições de solicitação. Educar em Revista, (62):69–85.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. In Advances in Neural Information Processing Systems, volume 36, pages 10088–10115.
Ferreira, R., Freitas, E., Cabral, L., Dawn, F., Rodrigues, L., Rakovic, M., Raniel, J., and Gasevic, D. (2024). Words of wisdom: A journey through the realm of nlp for learning analytics – a systematic literature review. Journal of Learning Analytics, 11(3):82–105.
Gemma Team (2025). Gemma 3 technical report. arXiv preprint arXiv:2503.19786.
Graham, S. and Harris, K. R. (2019). Evidence-based practices in writing. In Graham, S., MacArthur, C. A., and Hebert, M., editors, Best Practices in Writing Instruction. The Guilford Press, 3 edition.
Hou, Z., Ciuba, A., and Li, X. (2025). Improving llm-based automatic essay scoring with linguistic features. In Proceedings of the Innovation and Responsibility in AI-Supported Education Workshop, volume 273 of Proceedings of Machine Learning Research, pages 41–65. PMLR.
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022). Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
Hugging Face TB (2024). Smollm-135m model card. [link]. Accessed: 2026-05-30.
Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., and Kasneci, G. (2023). Chatgpt for good? on opportunities and challenges of large language models for education. Learning and Individual Differences, 103:102274.
Liew, P. Y. and Tan, I. K. T. (2024). On automated essay grading using large language models. In Proceedings of the 2024 8th International Conference on Computer Science and Artificial Intelligence.
Liu, Z., Zhao, C., Iandola, F., Lai, C., Tian, Y., Fedorov, I., Xiong, Y., Chang, E., Shi, Y., Krishnamoorthi, R., Lai, L., and Chandra, V. (2024). Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 32431–32454. PMLR.
Lobo, J., Anthony, L., Falcão, A., Xavier, C., Torrezão, N., Isotani, S., Pinto, I. I. B. S., Rodrigues, L. A. L., and Mello, R. F. (2025). Automatic scoring of elementary school essays in brazilian portuguese with llms: Comparing gemini, gpt-4o, claude, and mistral. In Anais do Simpósio Brasileiro de Informática na Educação.
Lu, Z., Li, X., Cai, D., Yi, R., Liu, F., Zhang, X., and Lane, N. D. (2024). Small language models: Survey, measurements, and insights. arXiv preprint arXiv:2409.15790.
Mello, R. F., Oliveira, H., Wenceslau, M., Batista, H., Cordeiro, T., Bittencourt, I. I., and Isotani, S. (2024). Propor'24 competition on automatic essay scoring of portuguese narrative essays. In Proceedings of the 16th International Conference on Computational Processing of Portuguese, pages 1–5.
Nguyen, C. V., Shen, X., Aponte, R., Xia, Y., Basu, S., Hu, Z., Chen, J., Parmar, M., Kunapuli, S., Barrow, J., Wu, J., Singh, A., Wang, Y., Gu, J., Ahmed, N. K., Lipka, N., Zhang, R., Chen, X., Yu, T., Kim, S., Deilamsalehy, H., Park, N., Rimer, M., Zhang, Z., Yang, H., Mathur, P., Wu, G., Dernoncourt, F., Rossi, R. A., and Nguyen, T. H. (2025). A survey on small language models. In Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing, pages 807–821. INCOMA Ltd.
Oliveira, H., Mello, R. F., Miranda, P., Batista, H., Silva Filho, M. W. d., Cordeiro, T., Pinto, I. I. B. S., and Isotani, S. (2025). A benchmark dataset of narrative student essays with multi-competency grades for automatic essay scoring in brazilian portuguese. Data in Brief, 60:111526.
Page, E. B. (1966). The imminence of grading essays by computer. The Phi Delta Kappan, 47(5):238–243.
Qwen Team (2024). Qwen2.5 technical report. arXiv preprint arXiv:2412.15115.
Seßler, K., Fürstenberg, M., Bühler, B., and Kasneci, E. (2025). Can ai grade your essays? a comparative analysis of large language models and teacher ratings in multidimensional essay scoring. In Proceedings of the 15th International Conference on Learning Analytics and Knowledge, pages 462–472. Association for Computing Machinery.
Shermis, M. D. and Burstein, J. (2013). Handbook of Automated Essay Evaluation: Current Applications and New Directions. Routledge.
Siemens, G. (2013). Learning analytics: The emergence of a discipline. American Behavioral Scientist, 57(10):1380–1400.
Silva Filho, M. W. d. and collaborators (2024). Brazilian portuguese narrative essays dataset. [link]. Accessed: 2026-05-30.
Silva Filho, M. W. d., Nascimento, A., Miranda, P., Rodrigues, L. A. L., Cordeiro, T., Isotani, S., Pinto, I. I. B. S., and Mello, R. F. (2023). Automated formal register scoring of student narrative essays written in portuguese. In Anais do II Workshop de Aplicações Práticas de Learning Analytics em Instituições de Ensino no Brasil, pages 1–11. SBC.
Slade, S. and Prinsloo, P. (2013). Learning analytics: Ethical issues and dilemmas. American Behavioral Scientist, 57(10):1510–1529.
Wang, D. and Wang, J. (2025). The impact mechanism of aes on improving english writing achievement. Scientific Reports, 15:3928.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. (2026). A survey of large language models. Frontiers of Computer Science, 20:2012627.
Publicado
05/10/2026
Como Citar
BATISTA, Hyan; ANTHONY, Lenon; FALCÃO, Andreza; ALVES, Gabriel; BATISTA, Maria da Conceicao Moraes; NASCIMENTO, André C. A.; MELLO, Rafael Ferreira.
SLMs como Avaliadores Automatizados: Avaliando Competências de Escrita Narrativa com Modelos Adaptados por LoRA. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 903-915.
DOI: https://doi.org/10.5753/sbie.2026.27403.
