Como Nossos Pais: Princípios Clássicos para a Avaliação de Sistemas de IA
Resumo
Este trabalho retoma bases epistemológicas e metodológicas de diversas áreas da ciência, como psicologia, linguística de corpus, sociologia e neurociência, para fundamentar princípios básicos da avaliação de sistemas baseados em Processamento de Linguagem Natural e Inteligência Artificial. Apresentamos em detalhe o estado da arte da avaliação na área, discutindo os motivos que nos levam à necessidade de automatizar a avaliação de sistemas (paradigma LLM-as-a-Judge). Introduzimos Cinco Princípios Metodológicos básicos que devem fundamentar a avaliação de sistemas baseados em IA, fundamentados tanto na teoria e experimentos destas áreas que já pesquisam a percepção humana há décadas quanto em evidências empíricas do próprio PLN.
Referências
Agarwal, R., Singh, A., Zhang, L. M., Bohnet, B., Chan, S., Anand, A., Abbas, Z., Nova, A., Co-Reyes, J. D., Chu, E., Behbahani, F., Faust, A., and Larochelle, H. (2024). Many-shot in-context learning. arXiv preprint arXiv:2404.11018.
Artemova, E., Tsvigun, A., Schlechtweg, D., Fedorova, N., Tilga, S., and Obmoroshev, B. (2025). Hands-on tutorial: Labeling with llm and human-in-the-loop. In Proceedings of COLING 2025.
Artstein, R. and Poesio, M. (2008). Inter-coder agreement for computational linguistics. Computational Linguistics, 34(4):555–596.
Asli, C., Clark, E., and Gao, J. (2020). Evaluation of text generation: A survey. arXiv:2006.14799.
Bayerl, P. S. and Paul, K. I. (2011). What determines inter-coder agreement in manual annotations? a meta-analytic investigation. Computational Linguistics, 37(4):699–725.
Beloff, H. (1958). Two forms of social conformity: Acquiescence and conventionality. The Journal of Abnormal and Social Psychology, 56(1):99–104.
Bird, S. and Liberman, M. (2001). A formal framework for linguistic annotation. Speech Communication, 33(1–2):23–60.
Blagec, K., Dorffner, G., Moradi, M., Ott, S., and Samwald, M. (2022). A global analysis of metrics used for measuring performance in natural language processing. In Proceedings of NLP Power! The First Workshop on Efficient Benchmarking in NLP, pages 52–63, Dublin, Ireland. Association for Computational Linguistics.
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D. (2015). A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), Lisbon, Portugal. Association for Computational Linguistics.
BRASIL. Conselho Nacional de Justiça (2025). Resolução n 615, de 10 de fevereiro de 2025. Conselho Nacional de Justiça, Brasília, DF.
Callison-Burch, C., Osborne, M., and Koehn, P. (2006). Re-evaluating the role of bleu in machine translation research. In Conference of the European Chapter of the Association for Computational Linguistics.
Carletta, J. (1996). Assessing agreement on classification tasks: The kappa statistic. Computational Linguistics, 22(2):249–254.
Chen, D., Chen, R., Zhang, S., Wang, Y., Liu, Y., Zhou, H., Zhang, Q., Wan, Y., Zhou, P., and Sun, L. (2024). MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark. In Forty-first International Conference on Machine Learning.
Chen, O., Castro-Alonso, J., Paas, F., and Sweller, J. (2018). Extending cognitive load theory to incorporate working memory resource depletion: Evidence from the spacing effect. Educational Psychology Review, 30(2):483–501.
Choi, J., Hong, Y., and Kim, B. (2025). People will agree what i think: Investigating LLM’s false consensus effect. In Findings of the Association for Computational Linguistics: NAACL 2025. Association for Computational Linguistics. arXiv:2407.12007v2.
Confident AI (2024). Deepeval: A framework for evaluating large language models. Accessed: 2026-05-12.
de Jong, T. (2010). Cognitive load theory, educational research, and instructional design: Some food for thought. Instructional Science, 38(2):105–134.
Durkin, K. (1996). Peer pressure. In Manstead, A. S. R. and Hewstone, M., editors, The Blackwell Encyclopedia of Social Psychology. Blackwell Publishers, Oxford.
Fanous, A., Goldberg, J., Agarwal, A. A., Lin, J., Zhou, A., Daneshjou, R., and Koyejo, S. (2025). Syceval: Evaluating llm sycophancy. arXiv preprint arXiv:2502.08177.
Fayyaz, M., Yin, F., Sun, J., and Peng, N. (2024). Evaluating human alignment and model faithfulness of llm rationale.
Foody, G. M. (2024). Ground truth in classification accuracy assessment: Myth and reality. Geomatics, 4(1):81–90.
Fort, K., Adda, G., and Cohen, K. B. (2011). Last words: Amazon mechanical turk: Gold mine or coal mine? Computational Linguistics, 37(2):413–420.
Freitas, C. (2024). Dataset e corpus. In Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, chapter 16. BPLN, 4 edition.
Frenda, S., Abercrombie, G., Basile, V., Pedrani, A., Panizzon, R., Cignarella, A. T., Marco, C., and Bernardi, D. (2024). Perspectivist approaches to natural language processing: a survey: Perspectivist approaches to natural language processing... Lang. Resour. Eval., 59(2):1719–1746.
Gao, M., Hu, X., Yin, X., Ruan, J., Pu, X., and Wan, X. (2025). LLM-based NLG evaluation: Current status and challenges. Computational Linguistics, 51:661–687.
Gilbert, D. T. (1991). How mental systems believe. American Psychologist, 46(2):107–119.
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E. (2022). Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection.
Jacovi, A. and Goldberg, Y. (2020). Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4198–4212, Online. Association for Computational Linguistics.
Jain, S., Ahmed, U. Z., Sahai, S., and Leong, B. (2025). Beyond consensus: Mitigating the agreeableness bias in llm judge evaluations.
Jang, M. E. and Silavong, F. (2025). InstaJudge: Aligning judgment bias of LLM-as-judge with humans in industry applications. In Potdar, S., Rojas-Barahona, L., and Montella, S., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 1158–1172, Suzhou (China). Association for Computational Linguistics.
Kalouli, A.-L., Buis, A., Real, L., Palmer, M., and de Paiva, V. (2019). Explaining simple natural language inference. In Proceedings of the 13th Linguistic Annotation Workshop, pages 132–143, Florence, Italy. Association for Computational Linguistics.
Kalouli, A.-L., Real, L., and de Paiva, V. (2017). Textual inference: getting logic from humans. In Gardent, C. and Retoré, C., editors, Proceedings of the 12th International Conference on Computational Semantics (IWCS) — Short Papers, Montpellier, France. Association for Computational Linguistics.
Laux, J., Stephany, F., and Liefgreen, A. (2024). Improving task instructions for data annotators: How clear rules and higher pay increase performance in data annotation in the ai economy.
Lee, Y., Kim, J., Kim, J., Cho, H., Kang, J., Kang, P., and Kim, N. (2025). Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), Suzhou, China. Association for Computational Linguistics.
Li, K., Li, Y., Zhang, T., Luo, H., Wu, X., Glass, J. R., and Meng, H. M. (2025a). Ragzeval: Enhancing rag response evaluation through end-to-end reasoning and ranking-based reinforcement learning. In Proceedings of EMNLP 2025.
Li, S., Li, J., Qu, Y., Shi, X., Guo, Y., He, Z., Wang, Y., and Tan, W. (2025b). Semantic-eval : A semantic comprehension evaluation framework for large language models generation without training. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T., editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9675–9690, Vienna, Austria. Association for Computational Linguistics.
Lin, S., Hilton, J., and Evans, O. (2021). Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.
Madsen, A. (2024). New faithfulness-centric interpretability paradigms for natural language processing. arXiv preprint arXiv:2411.17992.
Marasović, A. (2018). Nlp’s generalization problem, and how researchers are tackling it. The Gradient. Accessed: 2026-05-14.
Marelli, M., Menini, S., Baroni, M., Bentivogli, L., Bernardi, R., and Zamparelli, R. (2014). A SICK cure for the evaluation of compositional distributional semantic mo dels. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pages 216–223, Reykjavik, Iceland. European Language Resources Association (ELRA).
Marinho, M., Vianna, D., Migliorini, G., Real, L., Ramalhete, M., and da Silva, A. (2026). Grounding the evaluation of ai-generated legal answers: The hermeval protocol and empirical evidence. In Proceedings of the International Conference on Artificial Intelligence and Law (ICAIL), Singapore. ACM.
Martin, D., Hanrahan, B. V., O’Neill, J., and Gupta, N. (2014). Being a turker. In Proceedings of the 17th ACM Conference on Computer Supported Cooperative Work & Social Computing, CSCW ’14, page 224–235, New York, NY, USA. Association for Computing Machinery.
McEnery, T. and Hardie, A. (2012). Corpus Linguistics: Method, Theory and Practice. Cambridge University Press, Cambridge, 2 edition.
Messick, S. and Jackson, D. N. (1961). Acquiescence and the factorial interpretation of personality tests. In Jackson, D. N. and Messick, S., editors, Psychological Measurement in Personality Research, pages 61–103. University of Chicago Press, Chicago.
Meyer, C. F. (2002). English Corpus Linguistics: An Introduction. Cambridge University Press, Cambridge.
Novikova, J., Dušek, O., Cercas Curry, A., and Rieser, V. (2017). Why we need new evaluation metrics for NLG. In Palmer, M., Hwa, R., and Riedel, S., editors, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2241–2252, Copenhagen, Denmark. Association for Computational Linguistics.
Pagano, T. P., Loureiro, R. B., Lisboa, F. V. N., Peixoto, R. M., Guimarães, G. A. S., Cruz, G. O. R., Araujo, M. M., Santos, L. L., Cruz, M. A. S., Oliveira, E. L. S., Winkler, I., and Nascimento, E. G. S. (2023). Bias and unfairness in machine learning models: A systematic review on datasets, tools, fairness metrics, and identification and mitigation methods. Big Data and Cognitive Computing, 7(1).
Pavlick, E. and Kwiatkowski, T. (2019). Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677–694.
Plank, B. (2022). The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022), pages 10671–10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
Pradhan, V. K., Schaekermann, M., and Lease, M. (2021). In search of ambiguity: A three-stage workflow design to clarify annotation guidelines for crowd workers.
Qian, M., Sun, G., Gales, M. J. F., and Knill, K. M. (2026). Who can we trust? llm-as-a-jury for comparative assessment.
Raina, V., Liusie, A., and Gales, M. (2024). Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment. arXiv preprint arXiv:2402.14016.
Reiter, E. and Belz, A. (2009). An investigation into the validity of some metrics for automatically evaluating natural language generation systems. Computational Linguistics, 35(4):529–558.
Ross, L., Greene, D., and House, P. (1977). The “false consensus effect”: An egocentric bias in social perception and attribution processes. Journal of Experimental Social Psychology, 13(3):279–301.
Sandan, I. B., Dinh, T. A., and Niehues, J. (2025). Knockout LLM assessment: Using large language models for evaluations through iterative pairwise comparisons. In Arviv, O., Clinciu, M., Dhole, K., Dror, R., Gehrmann, S., Habba, E., Itzhak, I., Mille, S., Perlitz, Y., Santus, E., Sedoc, J., Shmueli Scheuer, M., Stanovsky, G., and Tafjord, O., editors, Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²), pages 121–128, Vienna, Austria and virtual meeting. Association for Computational Linguistics.
Snow, R., O’Connor, B., Jurafsky, D., and Ng, A. Y. (2008). Cheap and fast—but is it good? evaluating non-expert annotations for natural language tasks. In Proceedings of EMNLP 2008, Honolulu, Hawaii. Association for Computational Linguistics.
Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2):257–285.
Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D. (2024). Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges. arXiv preprint arXiv:2406.12624.
Trautmann, D., Ostapuk, N., Grail, Q., Pol, A. A., Bonifazi, G., Gao, S., and Gajek, M. (2024). Measuring the groundedness of legal question-answering systems. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP).
Tversky, A. and Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157):1124–1131.
Verga, P., Hofstätter, S., Althammer, S., Su, Y., Piktus, A., Arkhangorodsky, A., Xu, M., White, N., and Lewis, P. (2024). Replacing judges with juries: Evaluating llm generations with a panel of diverse models. arXiv preprint arXiv:2404.18796.
Voutilainen, A. (2012). Improving corpus annotation productivity: a method and experiment with interactive tagging. In Calzolari, N., Choukri, K., Declerck, T., Doğan, M. U., Maegaard, B., Mariani, J., Moreno, A., Odijk, J., and Piperidis, S., editors, Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 2097–2102, Istanbul, Turkey. European Language Resources Association (ELRA).
Wei, H., He, S., Xia, T., Liu, F., Wong, A., Lin, J., and Han, M. (2025). Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates.
Wojcieszak, M. and Price, V. (2009). What underlies the false consensus effect? how personal opinion and disagreement affect perception of public opinion. International Journal of Public Opinion Research, 21(1):25–46.
Wu, X., Xiao, L., Sun, Y., Zhang, J., Ma, T., and He, L. (2022). A survey of human-in-the-loop for machine learning. Future Generation Computer Systems, 135:364–381.
Yang, J., Redi, J., Demartini, G., and Bozzon, A. (2016). Modeling task complexity in crowdsourcing. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 4(1):249–258.
Zhao, W. X., Zhou, K., Li, J., Tang, T., and Wen, J.-R. (2026). Human Alignment, pages 209–256. Springer Nature Singapore, Singapore.
Zheng, L., Chiang, W.-L., Sheng, Y., and et al. (2023). Judging llm-as-a-judge with mt-bench and chatbot arena. In NeurIPS Workshop on Instruction Tuning and Instruction Following.
Zhu, Z., Tieleman, O., Bukhtiyarov, A., and Chen, J. (2026). Cyclicjudge: Mitigating judge bias efficiently in llm-based evaluation.
