Do LLM Judges Agree with Humans? Evaluating Financial Commentaries from Material Facts
Resumo
The rapid adoption of large language models (LLMs) in finance has enabled automated generation of in-domain texts such as earnings summaries, market analyses, and commentary on regulated disclosures. Generating accurate and accessible financial commentary from material facts poses challenges related to domain-specific language, strict factual faithfulness, and readability. Evaluating such outputs is difficult: traditional automatic metrics overlook financial correctness and adequacy, while human expert assessment is reliable but costly and subjective. This paper proposes a multidimensional evaluation protocol for generating financial commentary in Portuguese and investigates the use of LLMs as evaluators (“LLMs-as-judges”) in this high-stakes, low-resource setting. We systematically compare human expert judgments and LLM-based evaluation preferences, while also investigating complementary dimensions such as writing quality, factuality, usefulness, and simplicity. To the best of our knowledge, this is the first proposal of an evaluation protocol for this task in Portuguese. Furthermore, we analyze where LLM judgments align with or diverge from human assessments, providing practical insights and recommendations for evaluation methodologies in financial AI applications.Referências
Abonizio, H., Almeida, T. S., Laitz, T., Junior, R. M., Bonás, G. K., Nogueira, R., and Pires, R. (2024). Sabiá-3 technical report.
Anthropic (2026). Claude opus 4.6 system card. [link].
Assis, G., Dutra, H., Vianna, D., Meira, Wagner, J., da Silva, A. S., and Paes, A. (2025a). Language models for automated market commentary from corporate disclosures. In Proceedings of the 6th ACM International Conference on AI in Finance, ICAIF ’25, page 727–735, New York, NY, USA. Association for Computing Machinery.
Assis, G., Freitas, C., and Paes, A. (2025b). Exploring Brazil’s LLM Fauna: Investigating the generative performance of large language models in portuguese. Journal of the Brazilian Computer Society, 31(1):939–971.
Celikyilmaz, A., Clark, E., and Gao, J. (2020). Evaluation of text generation: A survey. arXiv preprint arXiv:2006.14799.
Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B. (2024). Humans or LLMs as the judge? a study on judgement bias. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8301–8327, Miami, Florida, USA. Association for Computational Linguistics.
Cohere Labs (2024). Introducing command r+: A scalable llm built for business. [link].
Deng, M., Tan, B., Liu, Z., Xing, E., and Hu, Z. (2021). Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t., editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Elo, A. E. (1978). The Rating of Chessplayers, Past and Present. Arco Publishing.
Gao, M., Hu, X., Yin, X., Ruan, J., Pu, X., and Wan, X. (2025). LLM-based NLG Evaluation: Current Status and Challenges. Computational Linguistics, 51(2):661–687.
Gemma Team (2025). Gemma 3 Technical Report.
Goldsack, T., Wang, Y., Lin, C., and Chen, C.-C. (2025). From facts to insights: A study on the generation and evaluation of analytical reports for deciphering earnings calls. In Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., and Schockaert, S., editors, Proceedings of the 31st International Conference on Computational Linguistics, pages 10576–10593, Abu Dhabi, UAE. Association for Computational Linguistics.
Google DeepMind (2026). Gemini 3.1 pro model card. [link]. Accessed: 2026.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M. A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2023). Mistral 7b.
Jin, S., Li, S., Zhang, S., and Yan, R. (2026). Finrpt: Dataset, evaluation system and LLM-based multi-agent framework for equity research report generation. Proceedings of the AAAI Conference on Artificial Intelligence, 40(1):507–515.
Lin, C.-Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. ACL.
Lipovetsky, S. and Conklin, M. (2003). Priority estimations by pair comparisons: Ahp, thurstone scaling, bradley-terry-luce, and markov stochastic modeling. In Proceedings of the ASA Joint Statistical Meeting, JSM.
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023). G-eval: NLG evaluation using gpt-4 with better human alignment. In Bouamor, H., Pino, J., and Bali, K., editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
Masid, M. R., Assis, G., Vianna, D., Paes, A., and Silva, A. S. d. (2026). Auditing the evaluators: How far can automatic evaluation go in assessing Portuguese financial texts? In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 222–233, Salvador, Brazil. Association for Computational Linguistics.
Meta AI (2024). The llama 3 herd of models.
Meta AI (2025). The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.
Mistral AI (2025). Mistral Small 3.1.
Mukherjee, R., Bohra, A., Banerjee, A., Sharma, S., Hegde, M., Shaikh, A., Shrivastava, S., Dasgupta, K., Ganguly, N., Ghosh, S., and Goyal, P. (2022). ECTSum: A new benchmark dataset for bullet point summarization of long earnings call transcripts. In Goldberg, Y., Kozareva, Z., and Zhang, Y., editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10893–10906, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
OpenAI (2024). Gpt-4o system card.
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on ACL, ACL ’02, page 311–318, USA. ACL.
Ribeiro, J. V. A., Correia, T. P., Requena, J. V. S. C., and Berton, L. (2026). Evaluating reference-free summarization quality metrics for Portuguese: A study with human judgments in financial news. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 899–907, Salvador, Brazil. Association for Computational Linguistics.
Sandan, I. B., Dinh, T. A., and Niehues, J. (2025). Knockout LLM assessment: Using large language models for evaluations through iterative pairwise comparisons. In Ar viv, O., Clinciu, M., Dhole, K., Dror, R., Gehrmann, S., Habba, E., Itzhak, I., Mille, S., Perlitz, Y., Santus, E., Sedoc, J., Shmueli Scheuer, M., Stanovsky, G., and Tafjord, O., editors, Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²), pages 121–128, Vienna, Austria and virtual meeting. Association for Computational Linguistics.
Sharma, N., Agarwal, N., and Sirts, K. (2026). Towards consistent detection of cognitive distortions: Llm-based annotation and dataset-agnostic evaluation. In Piperidis, S., Bel, N., van den Heuvel, H., Ide, N., Krek, S., and Toral, A., editors, Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 10866–10882, Palma, Mallorca, Spain. European Language Resources Association (ELRA).
Takayanagi, T., Goldsack, T., Izumi, K., Lin, C., Takamura, H., and Chen, C.-C. (2025). Earnings2Insights: Analyst report generation for investment guidance. In Chen, C.-C., Winata, G. I., Rawls, S., Das, A., Chen, H.-H., and Takamura, H., editors, Proceedings of The 10th Workshop on Financial Technology and Natural Language Processing, pages 246–251, Suzhou, China. Association for Computational Linguistics.
Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D. (2024). Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges. Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics.
Wataoka, K., Takahashi, T., and Ri, R. (2024). Self-preference bias in LLM-as-a-judge. In Neurips Safe Generative AI Workshop 2024.
Xiao, Y., Sun, E., Luo, D., and Wang, W. (2025). Tradingagents: Multi-agents LLM financial trading framework. In The First MARW: Multi-Agent AI in the Real World Workshop at AAAI 2025.
Xie, Q., Han, W., Chen, Z., Xiang, R., Zhang, X., He, Y., Xiao, M., Li, D., Dai, Y., Feng, D., et al. (2024). Finben: A holistic financial benchmark for large language models. Advances in Neural Information Processing Systems, 37:95716–95743.
Yan, S. (2022). Disentangled variational topic inference for topic-accurate financial report generation. In Chen, C.-C., Huang, H.-H., Takamura, H., and Chen, H.-H., editors, Proceedings of the Fourth Workshop on Financial Technology and Natural Language Processing (FinNLP), pages 18–24, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. (2025). Qwen3 technical report.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J., and Stoica, I. (2023). Judging llm-as-a-judge with mt-bench and chatbot arena. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S., editors, Advances in Neural Information Processing Systems, volume 36, pages 46595–46623. Curran Associates, Inc.
Zhu, D., Lappas, T., and Rachidi, T. (2023). Commentary generation for financial markets. Expert Systems with Applications, 211:118364.
Anthropic (2026). Claude opus 4.6 system card. [link].
Assis, G., Dutra, H., Vianna, D., Meira, Wagner, J., da Silva, A. S., and Paes, A. (2025a). Language models for automated market commentary from corporate disclosures. In Proceedings of the 6th ACM International Conference on AI in Finance, ICAIF ’25, page 727–735, New York, NY, USA. Association for Computing Machinery.
Assis, G., Freitas, C., and Paes, A. (2025b). Exploring Brazil’s LLM Fauna: Investigating the generative performance of large language models in portuguese. Journal of the Brazilian Computer Society, 31(1):939–971.
Celikyilmaz, A., Clark, E., and Gao, J. (2020). Evaluation of text generation: A survey. arXiv preprint arXiv:2006.14799.
Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B. (2024). Humans or LLMs as the judge? a study on judgement bias. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8301–8327, Miami, Florida, USA. Association for Computational Linguistics.
Cohere Labs (2024). Introducing command r+: A scalable llm built for business. [link].
Deng, M., Tan, B., Liu, Z., Xing, E., and Hu, Z. (2021). Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t., editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Elo, A. E. (1978). The Rating of Chessplayers, Past and Present. Arco Publishing.
Gao, M., Hu, X., Yin, X., Ruan, J., Pu, X., and Wan, X. (2025). LLM-based NLG Evaluation: Current Status and Challenges. Computational Linguistics, 51(2):661–687.
Gemma Team (2025). Gemma 3 Technical Report.
Goldsack, T., Wang, Y., Lin, C., and Chen, C.-C. (2025). From facts to insights: A study on the generation and evaluation of analytical reports for deciphering earnings calls. In Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., and Schockaert, S., editors, Proceedings of the 31st International Conference on Computational Linguistics, pages 10576–10593, Abu Dhabi, UAE. Association for Computational Linguistics.
Google DeepMind (2026). Gemini 3.1 pro model card. [link]. Accessed: 2026.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M. A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2023). Mistral 7b.
Jin, S., Li, S., Zhang, S., and Yan, R. (2026). Finrpt: Dataset, evaluation system and LLM-based multi-agent framework for equity research report generation. Proceedings of the AAAI Conference on Artificial Intelligence, 40(1):507–515.
Lin, C.-Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. ACL.
Lipovetsky, S. and Conklin, M. (2003). Priority estimations by pair comparisons: Ahp, thurstone scaling, bradley-terry-luce, and markov stochastic modeling. In Proceedings of the ASA Joint Statistical Meeting, JSM.
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023). G-eval: NLG evaluation using gpt-4 with better human alignment. In Bouamor, H., Pino, J., and Bali, K., editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
Masid, M. R., Assis, G., Vianna, D., Paes, A., and Silva, A. S. d. (2026). Auditing the evaluators: How far can automatic evaluation go in assessing Portuguese financial texts? In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 222–233, Salvador, Brazil. Association for Computational Linguistics.
Meta AI (2024). The llama 3 herd of models.
Meta AI (2025). The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.
Mistral AI (2025). Mistral Small 3.1.
Mukherjee, R., Bohra, A., Banerjee, A., Sharma, S., Hegde, M., Shaikh, A., Shrivastava, S., Dasgupta, K., Ganguly, N., Ghosh, S., and Goyal, P. (2022). ECTSum: A new benchmark dataset for bullet point summarization of long earnings call transcripts. In Goldberg, Y., Kozareva, Z., and Zhang, Y., editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10893–10906, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
OpenAI (2024). Gpt-4o system card.
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on ACL, ACL ’02, page 311–318, USA. ACL.
Ribeiro, J. V. A., Correia, T. P., Requena, J. V. S. C., and Berton, L. (2026). Evaluating reference-free summarization quality metrics for Portuguese: A study with human judgments in financial news. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 899–907, Salvador, Brazil. Association for Computational Linguistics.
Sandan, I. B., Dinh, T. A., and Niehues, J. (2025). Knockout LLM assessment: Using large language models for evaluations through iterative pairwise comparisons. In Ar viv, O., Clinciu, M., Dhole, K., Dror, R., Gehrmann, S., Habba, E., Itzhak, I., Mille, S., Perlitz, Y., Santus, E., Sedoc, J., Shmueli Scheuer, M., Stanovsky, G., and Tafjord, O., editors, Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²), pages 121–128, Vienna, Austria and virtual meeting. Association for Computational Linguistics.
Sharma, N., Agarwal, N., and Sirts, K. (2026). Towards consistent detection of cognitive distortions: Llm-based annotation and dataset-agnostic evaluation. In Piperidis, S., Bel, N., van den Heuvel, H., Ide, N., Krek, S., and Toral, A., editors, Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 10866–10882, Palma, Mallorca, Spain. European Language Resources Association (ELRA).
Takayanagi, T., Goldsack, T., Izumi, K., Lin, C., Takamura, H., and Chen, C.-C. (2025). Earnings2Insights: Analyst report generation for investment guidance. In Chen, C.-C., Winata, G. I., Rawls, S., Das, A., Chen, H.-H., and Takamura, H., editors, Proceedings of The 10th Workshop on Financial Technology and Natural Language Processing, pages 246–251, Suzhou, China. Association for Computational Linguistics.
Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D. (2024). Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges. Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics.
Wataoka, K., Takahashi, T., and Ri, R. (2024). Self-preference bias in LLM-as-a-judge. In Neurips Safe Generative AI Workshop 2024.
Xiao, Y., Sun, E., Luo, D., and Wang, W. (2025). Tradingagents: Multi-agents LLM financial trading framework. In The First MARW: Multi-Agent AI in the Real World Workshop at AAAI 2025.
Xie, Q., Han, W., Chen, Z., Xiang, R., Zhang, X., He, Y., Xiao, M., Li, D., Dai, Y., Feng, D., et al. (2024). Finben: A holistic financial benchmark for large language models. Advances in Neural Information Processing Systems, 37:95716–95743.
Yan, S. (2022). Disentangled variational topic inference for topic-accurate financial report generation. In Chen, C.-C., Huang, H.-H., Takamura, H., and Chen, H.-H., editors, Proceedings of the Fourth Workshop on Financial Technology and Natural Language Processing (FinNLP), pages 18–24, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. (2025). Qwen3 technical report.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J., and Stoica, I. (2023). Judging llm-as-a-judge with mt-bench and chatbot arena. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S., editors, Advances in Neural Information Processing Systems, volume 36, pages 46595–46623. Curran Associates, Inc.
Zhu, D., Lappas, T., and Rachidi, T. (2023). Commentary generation for financial markets. Expert Systems with Applications, 211:118364.
Publicado
19/10/2026
Como Citar
ASSIS, Gabriel; VIANNA, Daniela; REAL, Livy; MASID, Marina Ramalhete; NEPOMUCENO, João; ROTTSCHAEFER, Eduardo; SILVA, Altigran Soares da; PAES, Aline.
Do LLM Judges Agree with Humans? Evaluating Financial Commentaries from Material Facts. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 26-40.
DOI: https://doi.org/10.5753/stil.2026.26484.
