Evaluation of LLMs in Answering Questions: A Case Study on POSCOMP

  • Carlos H. P. Silveira UFOPA
  • Florindo R. S. Carreteiro UFOPA
  • Fernando A. Sousa UFOPA
  • Fabio M. F. Lobato UFOPA / USP

Resumo


Several studies have investigated Large Language Models’ (LLMs) ability to answer questions from national educational assessments. However, specialized domains such as databases remain underexplored. Furthermore, current literature lacks a comprehensive analysis of how multimodal capabilities affect accuracy and stability across multiple runs. This study aims to fill these gaps by evaluating the accuracy and stability of five modern LLMs in solving 25 database questions from POSCOMP, using API requests, with hyperparameter control and five independent runs per configuration. The models investigated were GPT-5 Nano, GPT-5.4, Grok 4.1 Fast, Gemini 3 Flash Preview, and Gemini 3.1 Pro Preview, considering both LaTeX and image inputs. Additionally, we applied non-parametric statistical tests to evaluate the significance of performance differences between these input modalities. The results show that Gemini 3.1 Pro Preview demonstrated the best overall performance, outperforming the other models in both accuracy and consistency across runs, while GPT-5 Nano and Grok 4.1 Fast exhibited significant performance degradation under image inputs. Finally, the findings reinforce the importance of API-based evaluations with multiple independent runs to analyze not only the accuracy but also the reliability of responses generated by LLMs in technical domains.

Palavras-chave: databases, large language models, multimodality, POSCOMP, stability

Referências

Almeida, F. and Caminha, C. Evaluation of entry-level open-source large language models for information extraction from digitized documents. In Anais do XII Symposium on Knowledge Discovery, Mining and Learning. Belém, Brasil, pp. 25–32, 2024.

Atil, B., Aykent, S., Chittams, A., Fu, L., Passonneau, R. J., Radcliffe, E., Rajagopal, G. R., Sloan, A., Tudrej, T., Ture, F., Wu, Z., Xu, L., and Baldwin, B. Non-determinism of "deterministic"llm settings. [link], 2025.

Carreteiro, F. R. S., Sousa, F. A. d., Marcacini, R. M., and Lobato, F. M. F. Avaliação do impacto de entradas multimodais em LLMs: um estudo de caso de respostas ao POSCOMP. In Workshop sobre Educação em Computação (WEI). Maceió, Brasil, pp. 678–689, 2025.

Carvalho, L., Junior, C. C., and Mendonça, N. Evaluating the accuracy and stability of frontier llms on enade computer science questions. In Anais da XXXV Brazilian Conference on Intelligent Systems. Fortaleza, Brasil, pp. 260–274, 2025.

Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. [link], 2021.

Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y. Holistic evaluation of language models. [link], 2022.

Martínez-Plumed, F., Contreras-Ochando, L., Ferri, C., Hernández-Orallo, J., Kull, M., Lachiche, N., Ramírez-Quintana, M. J., and Flach, P. Crisp-dm twenty years later: From data mining processes to data science trajectories. IEEE Transactions on Knowledge and Data Engineering 33 (8): 3048–3061, 2021.

Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J. J., and Ioannidis, J. P. A. A manifesto for reproducible science. Nature Human Behaviour 1 (1): 0021, 2017.

OpenAI. Gpt-4o system card. [link], 2024.

Romero, V., Assis, G., Carvalho, J., Mann, P., and Paes, A. Opportunities vs. risks: Exploring automatic annotation of financial polarity biases via large language models. In Anais do XIII Symposium on Knowledge Discovery, Mining and Learning. Fortaleza, Brasil, pp. 1–8, 2025.

Saldanha, M. S. and Digiampietri, L. A. Chatgpt and bard performance on the poscomp exam. In Proceedings of the 20th Brazilian Symposium on Information Systems. Juiz de Fora, Brazil, 2024.

Taschetto, L. and Fileto, R. Using retrieval-augmented generation to improve performance of large language models on the brazilian university admission exam. In Anais do XXXIX Simpósio Brasileiro de Bancos de Dados. Florianópolis, Brazil, pp. 799–805, 2024.

Viegas, C., Gheyi, R., and Ribeiro, M. Assessing the Capability of LLMs in Solving POSCOMP Questions. Journal of the Brazilian Computer Society 31 (1): 990–1003, 2025.

Wang, S., Xu, T., Li, H., Zhang, C., Liang, J., Tang, J., Yu, P. S., and Wen, Q. Large language models for education: A survey and outlook. IEEE Signal Processing Magazine 42 (6): 51–63, 2025.

Xiao, H., Zhou, F., Liu, X., Liu, T., Li, Z., Liu, X., and Huang, X. A comprehensive survey of large language models and multimodal large language models in medicine. Information Fusion vol. 117, pp. 102888, 2025.
Publicado
19/10/2026
SILVEIRA, Carlos H. P.; CARRETEIRO, Florindo R. S.; SOUSA, Fernando A.; LOBATO, Fabio M. F.. Evaluation of LLMs in Answering Questions: A Case Study on POSCOMP. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 97-104. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.31826.