How Do I Say It Best? What Prompt Optimisation Can Reveal About Optimal Prompt Formulation In LLM-based Automated Short Answer Scoring
Resumo
Large language models (LLMs) have shown promise for automated short-answer scoring (ASAS), particularly in cold-start settings where task-specific training data are unavailable. However, prior studies typically relied on a small number of manually designed prompts, making it unclear whether the observed performance reflects the capabilities of the model or just the quality of the prompt. In this paper, we investigate whether prompt optimisation improves ASAS performance and what it can reveal about effective prompt design. Using three open-weight LLMs, three task formulations, and English language benchmark datasets from SemEval-2013, we compare human-written prompts with prompts optimised using the ProTeGi framework. Results show that prompt optimisation yields consistent performance gains across 135 paired configurations. Linguistic analyses further reveal that optimised prompts contain more discourse markers and shorter noun phrases, and these properties are predictive of downstream scoring performance. The findings suggest that prompt optimisation can serve not only as a performance-enhancement technique, but also as a tool for identifying linguistic characteristics of effective prompts for educational NLP tasks.
Palavras-chave:
Automated Short Answer Scoring, Prompt Optimisation, Large Language Models
Referências
Apertus, P., Hernández-Cano, A., Hägele, A., Huang, A. H., Romanou, A., Solergibert, A.-J., Pasztor, B., Messmer, B., Garbaya, D., Ďurech, E. F., et al. (2025). Apertus: Democratizing open and compliant llms for global language environments. arXiv preprint arXiv:2509.14233.
Audibert, J.-Y. and Bubeck, S. (2010). Best arm identification in multi-armed bandits. In COLT - 23rd Conference on Learning Theory, pages 13–p.
Bai, X. and Stede, M. (2023). A survey of current machine learning approaches to student free-text evaluation for intelligent tutoring. International Journal of Artificial Intelligence in Education, 33(4):992–1030.
Bexte, M., Horbach, A., and Zesch, T. (2022). Similarity-based content scoring—how to make s-bert keep up with bert. In Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2022), pages 118–123. Association for Computational Linguistics.
Bexte, M. and Zesch, T. (2025). Is lunch free yet? overcoming the cold-start problem in supervised content scoring using zero-shot llm-generated training data. In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), pages 144–159. Association for Computational Linguistics.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
Brysbaert, M. and New, B. (2009). Moving beyond Kučera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for american english. Behavior research methods, 41(4):977–990.
Burrows, S., Gurevych, I., and Stein, B. (2015). The eras and trends of automatic short answer grading. International Journal of Artificial Intelligence in Education, 25(1):60–117.
Chang, L.-H. and Ginter, F. (2024). Automatic short answer grading for finnish with chatgpt. Proceedings of the AAAI Conference on Artificial Intelligence, 38(21):23173–23181.
Dada, I. D., Akinwale, A. T., Osinuga, I. A., Ogbu, H. N., and Tunde-Adeleke, T.-J. (2025). iAttention transformer: An inter-sentence attention mechanism for automated grading. Mathematics, 13(18):2991.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C., and Solorio, T., editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
Dzikovska, M. O., Nielsen, R. D., and Leacock, C. (2013). Semeval-2013 task 7: The joint student response analysis and 8th recognizing textual entailment challenge. In Proceedings of SemEval 2013, pages 263–274.
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rocktäschel, T. (2023). Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797.
Ferreira Mello, R., Pereira Junior, C., Rodrigues, L., Pereira, F. D., Cabral, L., Costa, N., Ramalho, G., and Gasevic, D. (2025). Automatic short answer grading in the llm era: Does gpt-4 with prompt engineering beat traditional models? In Proceedings of the 15th international learning analytics and knowledge conference, pages 93–103.
Galhardi, L., Barbosa, C. R., de Souza, R. C. T., and Brancher, J. D. (2018). Portuguese automatic short answer grading. In Brazilian Symposium on Computers in Education (Simpósio Brasileiro de Informática na Educação-SBIE), volume 29, page 1373.
Gombert, S., Fink, A., Giorgashvili, T., Jivet, I., Di Mitri, D., Yau, J., Frey, A., and Drachsler, H. (2024). From the automated assessment of student essay content to highly informative feedback: a case study. International Journal of Artificial Intelligence in Education, 34(4):1378–1416.
Gombert, S., Sun, Z., Zehner, F., Lossjew, J., Wyrwich, T., Czinczel, B. K., Bednorz, D., Bernholt, S., Neumann, K., Harms, U., et al. (2026). Report on the BEA 2026 shared task on rubric-based short answer scoring for german. In Proceedings of the 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026). Association for Computational Linguistics.
Honnibal, M., Montani, I., Van Landeghem, S., and Boyd, A. (2020). spacy: Industrial-strength natural language processing in python.
Jeanquartier, F., Jean-Quartier, C., Rieder, P., Misirlić, V., Pasero, C., Hohensinner, R., Müller, H., and Holzinger, A. (2026). Assessing the carbon footprint of language models: Towards sustainability in AI. Resources, Conservation and Recycling, 226:108670.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2023). Mistral 7b.
Jimenez, S., Becerra, C., and Gelbukh, A. (2013). SOFTCARDINALITY: Hierarchical text overlap for student response analysis. In Manandhar, S. and Yuret, D., editors, Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013), pages 280–284, Atlanta, Georgia, USA. Association for Computational Linguistics.
Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., and Kasneci, G. (2023). Chatgpt for good? on opportunities and challenges of large language models for education. Learning and Individual Differences, 103:102274.
Kaya, M. and Cicekli, I. (2024). A hybrid approach for automated short answer grading. IEEE Access, 12:96332–96341.
Klein, R., Kyrilov, A., and Tokman, M. (2011). Automated assessment of short free-text responses in computer science using latent semantic analysis. In Proceedings of the 16th annual joint conference on Innovation and technology in computer science education, page 158–162. ACM.
Kortemeyer, G. (2023). Performance of the pre-trained large language model gpt-4 on automated short answer grading. CoRR, abs/2309.09338.
Kreidler, C. (1998). Introducing English semantics. Routledge.
Leacock, C. and Chodorow, M. (2003). C-rater: Automated scoring of short-answer questions. Computers and the Humanities, 37(4):389–405.
Lee, G.-G., Latif, E., Wu, X., Liu, N., and Zhai, X. (2024). Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence, 6:100213.
Leidinger, A., van Rooij, R., and Shutova, E. (2023). The language of prompting: What linguistic properties make a prompt successful? In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9210–9232, Singapore. Association for Computational Linguistics.
Levy, O., Zesch, T., Dagan, I., and Gurevych, I. (2013). UKP-BIU: Similarity and entailment metrics for student response analysis. In Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013), pages 285–289, Atlanta, Georgia, USA. Association for Computational Linguistics.
Li, Z., Tomar, Y., and Passonneau, R. J. (2021). A semantic feature-wise transformation relation network for automatic short answer grading. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6030–6040, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Livingston, S. (2009). Constructed-response test questions: Why we use them; how we score them. Technical Report 11, ETS.
Padó, U., Eryilmaz, Y., and Kirschner, L. (2023). Short-answer grading for german: Addressing the challenges. International Journal of Artificial Intelligence in Education, 34(4):1321–1352.
Poulton, A. and Eliens, S. (2021). Explaining transformer-based models for automatic short answer grading. In Proceedings of ICDTE 2021, pages 110–116.
Pryzant, R., Iter, D., Li, J., Lee, Y., Zhu, C., and Zeng, M. (2023). Automatic prompt optimization with "gradient descent" and beam search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7957–7968, Singapore. Association for Computational Linguistics.
Ramnath, K., Zhou, K., Guan, S., Mishra, S. S., Qi, X., Shen, Z., Wang, S., Woo, S., Jeoung, S., Wang, Y., Wang, H., Ding, H., Lu, Y., Xu, Z., Zhou, Y., Srinivasan, B., Yan, Q., Chen, Y., Ding, H., Xu, P., and Cheong, L. L. (2025). A systematic survey of automatic prompt optimization techniques. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 33078–33110, Suzhou, China. Association for Computational Linguistics.
Saha, S., Dhamecha, T. I., Marvaniya, S., Sindhgatta, R., and Sengupta, B. (2018). Sentence level or token level features for automatic short answer grading?: Use both. In Artificial Intelligence in Education, volume 10947 of Lecture Notes in Computer Science, pages 503–517. Springer.
Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., and Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv preprint arXiv:2402.07927, 1.
Sung, C., Dhamecha, T. I., and Mukhi, N. (2019). Improving short answer grading using transformer-based pre-training. In Artificial Intelligence in Education, volume 11625 of Lecture Notes in Computer Science, pages 469–481. Springer.
Wang, G., Cheng, S., Zhan, X., Li, X., Song, S., and Liu, Y. (2023). Openchat: Advancing open-source language models with mixed-quality data. arXiv preprint arXiv:2309.11235.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., and Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
Zesch, T., Horbach, A., and Zehner, F. (2023). To score or not to score: Factors influencing performance and feasibility of automatic content scoring of text responses. Educational Measurement: Issues and Practice, 42(1):44–58.
Audibert, J.-Y. and Bubeck, S. (2010). Best arm identification in multi-armed bandits. In COLT - 23rd Conference on Learning Theory, pages 13–p.
Bai, X. and Stede, M. (2023). A survey of current machine learning approaches to student free-text evaluation for intelligent tutoring. International Journal of Artificial Intelligence in Education, 33(4):992–1030.
Bexte, M., Horbach, A., and Zesch, T. (2022). Similarity-based content scoring—how to make s-bert keep up with bert. In Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2022), pages 118–123. Association for Computational Linguistics.
Bexte, M. and Zesch, T. (2025). Is lunch free yet? overcoming the cold-start problem in supervised content scoring using zero-shot llm-generated training data. In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), pages 144–159. Association for Computational Linguistics.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
Brysbaert, M. and New, B. (2009). Moving beyond Kučera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for american english. Behavior research methods, 41(4):977–990.
Burrows, S., Gurevych, I., and Stein, B. (2015). The eras and trends of automatic short answer grading. International Journal of Artificial Intelligence in Education, 25(1):60–117.
Chang, L.-H. and Ginter, F. (2024). Automatic short answer grading for finnish with chatgpt. Proceedings of the AAAI Conference on Artificial Intelligence, 38(21):23173–23181.
Dada, I. D., Akinwale, A. T., Osinuga, I. A., Ogbu, H. N., and Tunde-Adeleke, T.-J. (2025). iAttention transformer: An inter-sentence attention mechanism for automated grading. Mathematics, 13(18):2991.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C., and Solorio, T., editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
Dzikovska, M. O., Nielsen, R. D., and Leacock, C. (2013). Semeval-2013 task 7: The joint student response analysis and 8th recognizing textual entailment challenge. In Proceedings of SemEval 2013, pages 263–274.
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rocktäschel, T. (2023). Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797.
Ferreira Mello, R., Pereira Junior, C., Rodrigues, L., Pereira, F. D., Cabral, L., Costa, N., Ramalho, G., and Gasevic, D. (2025). Automatic short answer grading in the llm era: Does gpt-4 with prompt engineering beat traditional models? In Proceedings of the 15th international learning analytics and knowledge conference, pages 93–103.
Galhardi, L., Barbosa, C. R., de Souza, R. C. T., and Brancher, J. D. (2018). Portuguese automatic short answer grading. In Brazilian Symposium on Computers in Education (Simpósio Brasileiro de Informática na Educação-SBIE), volume 29, page 1373.
Gombert, S., Fink, A., Giorgashvili, T., Jivet, I., Di Mitri, D., Yau, J., Frey, A., and Drachsler, H. (2024). From the automated assessment of student essay content to highly informative feedback: a case study. International Journal of Artificial Intelligence in Education, 34(4):1378–1416.
Gombert, S., Sun, Z., Zehner, F., Lossjew, J., Wyrwich, T., Czinczel, B. K., Bednorz, D., Bernholt, S., Neumann, K., Harms, U., et al. (2026). Report on the BEA 2026 shared task on rubric-based short answer scoring for german. In Proceedings of the 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026). Association for Computational Linguistics.
Honnibal, M., Montani, I., Van Landeghem, S., and Boyd, A. (2020). spacy: Industrial-strength natural language processing in python.
Jeanquartier, F., Jean-Quartier, C., Rieder, P., Misirlić, V., Pasero, C., Hohensinner, R., Müller, H., and Holzinger, A. (2026). Assessing the carbon footprint of language models: Towards sustainability in AI. Resources, Conservation and Recycling, 226:108670.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2023). Mistral 7b.
Jimenez, S., Becerra, C., and Gelbukh, A. (2013). SOFTCARDINALITY: Hierarchical text overlap for student response analysis. In Manandhar, S. and Yuret, D., editors, Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013), pages 280–284, Atlanta, Georgia, USA. Association for Computational Linguistics.
Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., and Kasneci, G. (2023). Chatgpt for good? on opportunities and challenges of large language models for education. Learning and Individual Differences, 103:102274.
Kaya, M. and Cicekli, I. (2024). A hybrid approach for automated short answer grading. IEEE Access, 12:96332–96341.
Klein, R., Kyrilov, A., and Tokman, M. (2011). Automated assessment of short free-text responses in computer science using latent semantic analysis. In Proceedings of the 16th annual joint conference on Innovation and technology in computer science education, page 158–162. ACM.
Kortemeyer, G. (2023). Performance of the pre-trained large language model gpt-4 on automated short answer grading. CoRR, abs/2309.09338.
Kreidler, C. (1998). Introducing English semantics. Routledge.
Leacock, C. and Chodorow, M. (2003). C-rater: Automated scoring of short-answer questions. Computers and the Humanities, 37(4):389–405.
Lee, G.-G., Latif, E., Wu, X., Liu, N., and Zhai, X. (2024). Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence, 6:100213.
Leidinger, A., van Rooij, R., and Shutova, E. (2023). The language of prompting: What linguistic properties make a prompt successful? In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9210–9232, Singapore. Association for Computational Linguistics.
Levy, O., Zesch, T., Dagan, I., and Gurevych, I. (2013). UKP-BIU: Similarity and entailment metrics for student response analysis. In Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013), pages 285–289, Atlanta, Georgia, USA. Association for Computational Linguistics.
Li, Z., Tomar, Y., and Passonneau, R. J. (2021). A semantic feature-wise transformation relation network for automatic short answer grading. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6030–6040, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Livingston, S. (2009). Constructed-response test questions: Why we use them; how we score them. Technical Report 11, ETS.
Padó, U., Eryilmaz, Y., and Kirschner, L. (2023). Short-answer grading for german: Addressing the challenges. International Journal of Artificial Intelligence in Education, 34(4):1321–1352.
Poulton, A. and Eliens, S. (2021). Explaining transformer-based models for automatic short answer grading. In Proceedings of ICDTE 2021, pages 110–116.
Pryzant, R., Iter, D., Li, J., Lee, Y., Zhu, C., and Zeng, M. (2023). Automatic prompt optimization with "gradient descent" and beam search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7957–7968, Singapore. Association for Computational Linguistics.
Ramnath, K., Zhou, K., Guan, S., Mishra, S. S., Qi, X., Shen, Z., Wang, S., Woo, S., Jeoung, S., Wang, Y., Wang, H., Ding, H., Lu, Y., Xu, Z., Zhou, Y., Srinivasan, B., Yan, Q., Chen, Y., Ding, H., Xu, P., and Cheong, L. L. (2025). A systematic survey of automatic prompt optimization techniques. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 33078–33110, Suzhou, China. Association for Computational Linguistics.
Saha, S., Dhamecha, T. I., Marvaniya, S., Sindhgatta, R., and Sengupta, B. (2018). Sentence level or token level features for automatic short answer grading?: Use both. In Artificial Intelligence in Education, volume 10947 of Lecture Notes in Computer Science, pages 503–517. Springer.
Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., and Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv preprint arXiv:2402.07927, 1.
Sung, C., Dhamecha, T. I., and Mukhi, N. (2019). Improving short answer grading using transformer-based pre-training. In Artificial Intelligence in Education, volume 11625 of Lecture Notes in Computer Science, pages 469–481. Springer.
Wang, G., Cheng, S., Zhan, X., Li, X., Song, S., and Liu, Y. (2023). Openchat: Advancing open-source language models with mixed-quality data. arXiv preprint arXiv:2309.11235.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., and Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
Zesch, T., Horbach, A., and Zehner, F. (2023). To score or not to score: Factors influencing performance and feasibility of automatic content scoring of text responses. Educational Measurement: Issues and Practice, 42(1):44–58.
Publicado
05/10/2026
Como Citar
GOMBERT, Sebastian; HAHN, Sonja; ZEHNER, Fabian; CONG, Longwei; SUN, Zhifan; CAMUS, Leon; RIBEIRO, Fabíola Gonçalves C.; DRACHSLER, Hendrik.
How Do I Say It Best? What Prompt Optimisation Can Reveal About Optimal Prompt Formulation In LLM-based Automated Short Answer Scoring. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 1659-1671.
DOI: https://doi.org/10.5753/sbie.2026.28041.
