Comparing the Ability of LLMs to Identify Surmise Relations Between Competencies

Resumo


Adaptive instruction tailored to learners' knowledge states requires fine-grained assessment that can be interpreted regarding the next feasible learning steps. Modern psychometric frameworks leverage graphs of competencies connected by surmise relations. However, both data- and expert-based elicitation of these surmise relations scale only to a limited extent. This contribution investigates whether Large Language Models (LLMs) can serve as a scalable alternative by evaluating their performance in a zero-shot setting. Therefore, the results reveal a clear performance gap between commercial and open-source models, with Claude Haiku 4.5 and Gemini 2.5 flash achieving the best results. The findings suggest that top-tier LLMs can partially automate the elaboration of surmise relations, though human-in-the-loop validation remains necessary, particularly for semantically ambiguous relation types.
Palavras-chave: Automated curriculum analysis, Learning progressions, Prerequisite relations, Competence Structures

Referências

Anselmi, P., de Chiusole, D., and Heller, J. (2024). Adaptive assessment. In Heller, J. and Stefanutti, L., editors, Knowledge Structures: Recent Developments in Theory and Application, volume 7 of Advanced Series on Mathematical Psychology, pages 159–180. World Scientific.

Anthropic (2025). Introducing Claude Haiku 4.5. [link]. Accessed: June 1, 2026.

Bez, S., Burkart, F., Tomasik, M. J., and Merk, S. (2025). How do teachers process technology-based formative assessment results in their daily practice? results from process mining of think-aloud data. Learning and Instruction, 97:102100.

Chase, H. (2023). Langchain. [link].

Cosyn, E. and Thiéry, N. (2000). A practical procedure to build a knowledge structure. Journal of Mathematical Psychology, 44(3):383–407.

Cosyn, E., Uzun, H., Doble, C., and Matayoshi, J. (2021). A practical perspective on knowledge space theory: ALEKS and its data. Journal of Mathematical Psychology, 101:102512.

Daro, P., Mosher, F. A., and Corcoran, T. B. (2011). Learning trajectories in mathematics: A foundation for standards, curriculum, assessment, and instruction.

de Ayala, R. J. (2022). The Theory and Practice of Item Response Theory. Guilford Press, New York, 2nd edition.

de Chiusole, D., Spoto, A., and Stefanutti, L. (2024). Innovative methods for building knowledge structures. In Heller, J. and Stefanutti, L., editors, Knowledge Structures: Recent Developments in Theory and Application, volume 7 of Advanced Series on Mathematical Psychology, pages 87–104. World Scientific.

Doignon, J.-P. and Falmagne, J.-C. (1985). Spaces for the assessment of knowledge. International Journal of Man-Machine Studies, 23(2):175–196.

Google (2026). Gemma 3 4b instruction-tuned (Gemma-3-4b-it) model repository. [link]. Accessed: June 1, 2026.

Google Cloud (2026). Gemini 2.5 flash model documentation. [link]. Accessed: June 1, 2026.

Huff, K. and Goodman, D. P. (2007). The demand for cognitive diagnostic assessment. In Leighton, J. P. and Gierl, M. J., editors, Cognitive Diagnostic Assessment for Education: Theory and Applications, pages 19–60. Cambridge University Press, Cambridge.

Ihichr, A., Oustous, O., El Idrissi, Y. E. B., and Ait Lahcen, A. (2024). A systematic review on assessment in adaptive learning: Theories, algorithms and techniques. International Journal of Advanced Computer Science and Applications, 15(7).

Kambouri, M., Koppen, M., Villano, M., and Falmagne, J.-C. (1994). Knowledge assessment: tapping human expertise by the QUERY routine. International Journal of Human-Computer Studies, 40(1):119–151.

Käser, T., Klingler, S., Schwing, A. G., and Gross, M. (2017). Dynamic Bayesian networks for student modeling. IEEE Transactions on Learning Technologies, 10(4):450–462.

Koppen, M. and Doignon, J.-P. (1990). How to build a knowledge space by querying an expert. Journal of Mathematical Psychology, 34(3):311–331.

Le, N. L. and Abel, M.-H. (2025). How well do LLMs predict prerequisite skills? Zero-Shot comparison to expert-defined concepts. In Proceedings of the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pages 7245–7250.

Li, G., Wang, P., and Ke, W. (2023). Revisiting large language models as zero-shot relation extractors. pages 6877–6892.

NCTM, À. (2000). National council of teachers of mathematics (2000) principles and standards for school mathematics. Council of Teachers of Mathematics.

Ocheja, P., Flanagan, B., Dai, Y., and Ogata, H. (2024). How good is chatgpt in giving adaptive guidance using knowledge graphs in e-learning environments? arXiv preprint arXiv:2412.03856.

of Chief State School Officers, C. (2010). Common core state standards for mathematics. Accessed: 2026-06-01.

OpenAI (2026). Gpt-5-mini model documentation. [link]. Accessed: June 1, 2026.

Opitz, J. (2024). A closer look at classification evaluation metrics and a critical reflection of common evaluation practice. transactions of the association for computational linguistics, 12:820–836.

Pan, L., Li, C., Li, J., and Tang, J. (2017). Prerequisite relation learning for concepts in MOOCs. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1447–1456, Vancouver, Canada. Association for Computational Linguistics.

Qwen (2024). Qwen2.5-1.5B model repository. [link]. Accessed: June 1, 2026.

Roy, S., Madhyastha, M., Lawrence, S., and Rajan, V. (2019). Inferring concept prerequisite relations from online educational resources. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 9589–9594.

Sargin, A. and Ünlü, A. (2009). Inductive item tree analysis: Corrections, improvements, and comparisons. Mathematical Social Sciences, 58(3):376–392.

Segedinac, M. T., Horvat, S., Rodić, D. D., Rončević, T. N., and Savić, G. (2018). Using knowledge space theory to compare expected and real knowledge spaces in learning stoichiometry. Chemistry Education Research and Practice, 19(3):670–680.

Singh, A., Tan, D., Hepworth, C., and Seeland, M. (2024). Comparing LLM responses for education across model families. arXiv preprint arXiv:2502.08450.

Son, M. and Park, Y. (2025). Difficulties and solutions experienced by content experts in the process of developing a science curriculum ontology. Brain, Digital, Learning.

Templin, J. and Bradshaw, L. (2014). Hierarchical diagnostic classification models: A family of models for estimating and testing attribute hierarchies. Psychometrika, 79(2):317–339.

Thompson, W. J. and Clark, A. K. (2024). Improving instructional decision-making using diagnostic classification models. Educational Measurement: Issues and Practice.

Thompson, W. J. and Nash, B. (2022). A diagnostic framework for the empirical evaluation of learning maps. Frontiers in Education, 6:714736.

Ünlü, A. and Schrepp, M. (2021). Generalized inductive item tree analysis. Journal of Mathematical Psychology, 103:102547.

Wu, X., Sun, S., Xu, T., and Wang, A. (2024). Research on the selection of cognitive diagnosis model from the perspective of experts. Current Psychology, 43:13802–13810.

Zhang, M., Wang, J., Xiao, K., Wang, S., Zhang, Y., Chen, H., and Li, Z. (2025). Learning concept prerequisite relation via global knowledge relation optimization. Proceedings of the AAAI Conference on Artificial Intelligence, 39(2):1638–1646.
Publicado
05/10/2026
STEINER, Peter; LOBO, Jamilla Soares; MELLO, Rafael. Comparing the Ability of LLMs to Identify Surmise Relations Between Competencies. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 446-459. DOI: https://doi.org/10.5753/sbie.2026.27194.