Compression-Based Linguistic Complexity Metrics in Automatic Essay Scoring
Resumo
Compression-based linguistic complexity metrics enable cross-linguistic comparison without prior annotation. Their sensitivity to variation across languages and Portuguese registers highlights their applicability in NLP tasks. This study investigates their use as readability proxies and complementary features in Automatic Essay Scoring. We analyze how these metrics capture variation in essay quality across traits, genres, and educational levels in Brazilian Portuguese. In addition, we evaluate their sensitivity to differences between humanand AI-generated essays. Our results suggest that complexity metrics are effective (i) in differentiating educational levels, (ii) in detecting whether they were written by humans and (iii) as predictors of essay quality.
Referências
Amorim, E. C. F. and Veloso, A. (2017). A multi-aspect analysis of automatic essay scoring for Brazilian Portuguese. In Proceedings of the Student Research Workshop at the 15th Conference of the European Chapter of the Association for Computational Linguistics, pages 94–102. Association for Computational Linguistics.
Barbosa, A., Silveira, I. C., and Mauá, D. D. (2025). An empirical analysis of large language models for automated cross-prompt essay trait scoring in brazilian portuguese. Journal of the Brazilian Computer Society, 31(1):857–870.
Bazelato, B. S. and Amorim, E. C. F. (2013). A bayesian classifier to automatic correction of portuguese essays. In Conferência Internacional sobre Informática na Educação (TISE), volume 18, pages 779–782.
Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1):289–300.
Biber, D. (1995). Dimensions of Register Variation: A Cross-Linguistic Comparison. Cambridge University Press, 1st edition.
Biber, D. and Conrad, S. (2009). Register, Genre, and Style. Cambridge University Press, 1st edition.
Boquio, E. N. V. and Naval, Jr., P. C. (2024). Beyond canonical fine-tuning: Leveraging hybrid multi-layer pooled representations of BERT for automated essay scoring. In Calzolari, N., Kan, M.-Y., Hoste, V., Lenci, A., Sakti, S., and Xue, N., editors, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 2285–2295, Torino, Italia. ELRA and ICCL.
Brideau, L. (2025). Correlation: Pearson, spearman, and kendall’s tau. UVA Library StatLab, University of Virginia.
Cavalcanti, R., Casini, G., Assis, G., Real, L., Vianna, D., Mann, P., and Paes, A. (2025). Diplomatrix-br: Um corpus paralelo de redaçoes de autoria humana e de llms no concurso de diplomacia brasileira. In Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL), pages 192–205. SBC.
Chalegre, P. C., Machado, V. d. R., and Feltrim, V. D. (2026). Avaliação automática de redações do enem: Uma análise comparativa entre engenharia de características e transformers. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 738–748, Salvador, Brazil. Association for Computational Linguistics.
Croft, W. (2002). Typology and Universals. Cambridge University Press, 2 edition.
Ehret, K. (2021). An information-theoretic view on language complexity and register variation: Compressing naturalistic corpus data. Corpus Linguistics and Linguistic Theory, 17(2):383–410.
Ehret, K. and Szmrecsanyi, B. (2016). An information-theoretic approach to assess linguistic complexity, pages 71–94. De Gruyter, Berlin, Boston.
Fonseca, E. R., Medeiros, I., Kamikawachi, D., and Bokan, A. (2018). Automatically grading brazilian student essays. In Computational Processing of the Portuguese Language. PROPOR 2018., pages 170–179.
Grünwald, P. D. and Vitányi, P. M. (2008). Algorithmic information theory. In Adriaans, P. and van Benthem, J., editors, Philosophy of Information, Handbook of the Philosophy of Science, pages 281–317. North-Holland, Amsterdam.
Hockett, C. (1958). A Course in Modern Linguistics. A Course in Modern Linguistics. Macmillan.
Hollander, M., Wolfe, D. A., and Chicken, E. (2014). Nonparametric Statistical Methods. John Wiley & Sons, Inc., Hoboken, New Jersey, 3rd edition.
Juola, P. (1998). Measuring linguistic complexity: The morphological tier. Journal of Quantitative Linguistics, 5(3):206–213.
Juola, P. (2008). Assessing linguistic complexity, pages 89–108. John Benjamins Publishing Company.
Kendall, M. G. (1945). The treatment of ties in ranking problems. Biometrika, 33(3):239–251.
Laerd Statistics (2018). Kendall’s tau-b using spss statistics - a how-to statistical guide.
Leal, S. E., Duran, M. S., Scarton, C. E., Hartmann, N. S., and Aluísio, S. M. (2024). NILC-Metrix: Assessing the complexity of written and spoken language in Brazilian Portuguese. Language Resources and Evaluation, 58(1):73–110.
Leal, S. E., Serras, F. R., Finger, M., and Aluísio, S. M. (2026). Complexidade textual e suas tarefas relacionadas. In Caseli, H. M. and Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, volume 3, book chapter 6. BPLN, 4 edition.
Liu, H. (2015). Comparing Welch’s ANOVA, a Kruskal-Wallis test and traditional ANOVA in case of heterogeneity of variance. Master’s thesis, Virginia Commonwealth University, Richmond, VA. VCU Theses and Dissertations.
Liu, J., Xu, Y., and Zhu, Y. (2019). Automated essay scoring based on two-stage learning.
Liu, Y., Zhang, Z., Zhang, W., Yue, S., Zhao, X., Cheng, X., Zhang, Y., and Hu, H. (2023). Argugpt: evaluating, understanding and identifying argumentative essays generated by gpt models.
Lobo, T. R. and Martins, C. A. (2026). Evolução de padrões linguísticos na escrita científica em português: Uma análise com NILC-metrix. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 1062–1067, Salvador, Brazil. Association for Computational Linguistics.
Locatelli, M. S., Miranda, M. P., Da Silva Costa, I. J., Prates, M. T., Thomé, V., Zaparoli Monteiro, M., Lacerda, T., Pagano, A., Rios Neto, E., Meira Jr., W., and Almeida, V. (2024). Examining the behavior of llm architectures within the framework of standardized national exams in brazil. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 7:879–890.
Marinho, J. C., Anchiêta, R. T., and Moura, R. S. (2021). Essay-br: a brazilian corpus of essays. In Anais do III Dataset Showcase Workshop, pages 53–64. Sociedade Brasileira de Computação.
Marinho, J. C., Cordeiro, F., Anchiêta, R. T., and Moura, R. S. (2022). Automated essay scoring: An approach based on enem competencies. In Anais do XIX Encontro Nacional de Inteligência Artificial e Computacional, pages 49–60. SBC.
Matos, G. G. d. and Feltrim, V. D. (2026). Avaliação automática de redações do enem: Um estudo empírico sobre representações linguísticas e contextuais. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 488–497, Salvador, Brazil. Association for Computational Linguistics.
McNamara, D. S., Graesser, A. C., McCarthy, P. M., and Cai, Z. (2014). Automated Evaluation of Text and Discourse with Coh-Metrix. Cambridge University Press, 1st edition.
McWhorter, J. H. (2001). The worlds simplest grammars are creole grammars. Linguistic Typology, 5(2-3).
Mello, R. F., Oliveira, H., Wenceslau, M., Batista, H., Cordeiro, T., Bittencourt, I. I., and Isotanif, S. (2024). Propor’24 competition on automatic essay scoring of portuguese narrative essays. In Proceedings of the 16th International Conference on Computational Processing of Portuguese-Vol. 2, pages 1–5.
Monteiro, R., Correia, S., Amaro, R., Moutinho, M., Barbosa, S., and Reis, M. L. (2025). Níveis e descritores de complexidade textual para adultos de baixa literacia: um referencial do projeto iread4skills. Revista da Associação Portuguesa de Linguística, (13):193–222.
Nichols, J. (1998). Linguistic Diversity in Space and Time. University of Chicago Press.
Page, E. B. (1966). The imminence of... grading essays by computer. The Phi Delta Kappan, pages 238–243.
Rossman, L. N., Silveira, I. C., and Mauá, D. D. (2026). Evaluating automated scoring models on official ENEM essays. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 161–171, Salvador, Brazil. Association for Computational Linguistics.
Sardinha, T. B., Kauffmann, C., and Acunzo, C. M. (2014). A multi-dimensional analysis of register variation in Brazilian Portuguese. Corpora, 9(2):239–271.
Serras, F., Carpi, M., Branco, M., and Finger, M. (2024). Analysing and validating language complexity metrics across South American indigenous languages. In Kuribayashi, T., Rambelli, G., Takmaz, E., Wicke, P., and Oseki, Y., editors, Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 152–165, Bangkok, Thailand. Association for Computational Linguistics.
Serras, F. R., Carpi, M. D. M., Sturzeneker, M. L., Palma, M. F., Costa, A. S., Monte, V. M. D., Namiuti, C., Crespo, M. C. R. M., Paixão De Sousa, M. C., and Finger, M. (2026). Análise e Classificação Automática de Domínios Discursivos no Português do Brasil. Linguamática, 17(2):131–171.
Serras, F. R. and Finger, M. (2026). Compression-based language complexity under register variation in Portuguese. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 808–818, Salvador, Brazil. Association for Computational Linguistics.
Silveira, I. C., Barbosa, A., da Costa, D. S. L., and Mauá, D. D. (2025a). Investigating universal adversarial attacks against transformers-based automatic essay scoring systems. In Paes, A. and Verri, F. A. N., editors, Intelligent Systems, pages 169–183, Cham. Springer Nature Switzerland.
Silveira, I. C., Barbosa, A., and Mauá, D. D. (2024). A new benchmark for automatic essay scoring in Portuguese. In Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1, pages 228–237.
Silveira, I. C. and Mauá, D. D. (2026). Neuro-symbolic approaches for rubric-based automatic essay evaluation of ENEM essays. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 790–799, Salvador, Brazil. Association for Computational Linguistics.
Silveira, I. C., Ribeiro, E., Mamede, N., and Baptista, J. (2025b). Aprendizado por transferência para correçao automática de redaçao. Linguamática, 17(2):99–116.
Welch, B. L. (1951). On the comparison of several mean values: An alternative approach. Biometrika, 38(3/4):330–336.
