Tail Smoothing and Cross-Lingual Volatility: Evaluating Estimative Uncertainty in Large Language Models for Brazilian Portuguese

  • Renato O. Miyaji USP
  • Pedro L. P. Corrêa USP

Resumo


This study evaluates how LLMs interpret Words of Estimative Probability (WEPs) in Brazilian Portuguese compared to English. We translated an English benchmark and compared multilingual models (GPT-5.1, Gemini 3 Flash) against a region-specific model (Sabiá 4). Our findings reveal a “tail smoothing” phenomenon, where models systematically compress extreme probabilities. Notably, while Gemini 3 Flash demonstrated remarkable cross-lingual stability, GPT-5.1 exhibited significant calibration degradation. Counterintuitively, when measured against the English human baseline, Sabiá 4 displayed the most aggressive distribution compression. These results suggest a complex dynamic: either linguistic fine-tuning does not ensure alignment, or it actively captures a culturally specific pragmatic ambiguity in Brazilian Portuguese. This exposes the vulnerabilities and nuances of deploying deploying LLMs in nuanced semantic tasks in Brazilian Portuguese.

Referências

Erev, I. and Cohen, B. L. (1990). Verbal versus numerical probabilities: efficiency, biases, and the preference paradox. Organizational Behavior and Human Decision Processes, 45:1–18.

Fagen-Ulmschneider, W. (2019). Perception of probability words. University of Illinois at Urbana-Champaign.

Freitag, R. M. K., Cardoso, P. B., and Tejada, J. (2022). Linguistic and paralinguistic constraints on the function of (eu) acho que as discourse markers in brazilian portuguese: A multilevel approach. Pragmatics & Cognition, 29(2):324–346.

Google DeepMind (2026). Gemini 3 flash. Technical report. Multimodal large language model developed by Google DeepMind.

Juanchich, M. and Sirota, M. (2020). Do people really prefer verbal probabilities? Psychological Research, 84:2325–2338.

Laitz, T., Almeida, T. S., Abonizio, H., Malaquias Junior, R., Bonás, G. K., Piau, M., Larcher, C., Pires, R., and Nogueira, R. (2026). Sabiá-4 technical report. arXiv preprint arXiv:2603.10213.

OpenAI (2026). Gpt-5.1. Technical report. Large language model developed by OpenAI.

Schneider, S. (2016). Communicating uncertainty: A challenge for science communication. In Drake, J., Kontar, Y., Eichelberger, J., Rupp, T., and Taylor, K., editors, Communicating Climate-Change and Natural Hazard Risk and Cultivating Resilience, volume 45 of Advances in Natural and Technological Hazards Research. Springer, Cham.

Schuck, A. d. F., Garcia, G. L., Manesco, J. R. R., Paiola, P. H., et al. (2025). Evaluating large language models for brazilian portuguese sentiment analysis: A comparative study of multilingual state-of-the-art vs. brazilian portuguese fine-tuned llms. Journal of the Brazilian Computer Society, 31(1):885–917.

Shorinwa, O., Mei, Z., Lidard, J., Ren, A. Z., and Majumdar, A. (2025). A survey on uncertainty quantification of large language models: Taxonomy, open research challenges, and future directions. ACM Computing Surveys, 58(3).

Tang, Z., Shen, K., and Kejriwal, M. (2026). An evaluation of estimative uncertainty in large language models. npj Complex, 3:8.

Vlasyan, G. R. et al. (2018). Linguistic hedging in the light of politeness theory. In European Proceedings of Social and Behavioural Sciences. Future Academy.

Yona, G., Aharoni, R., and Geva, M. (2024). Can large language models faithfully express their intrinsic uncertainty in words? In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 7752–7764, Miami, Florida, USA. Association for Computational Linguistics.
Publicado
19/10/2026
MIYAJI, Renato O.; CORRÊA, Pedro L. P.. Tail Smoothing and Cross-Lingual Volatility: Evaluating Estimative Uncertainty in Large Language Models for Brazilian Portuguese. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 258-269. DOI: https://doi.org/10.5753/stil.2026.24563.