SPEAR: Um Método para Descoberta de Padrões Psicolinguísticos e Identificação de Discursos Presidenciais Destoantes
Resumo
A análise computacional de textos tem ampliado a investigação de como escolhas linguísticas revelam traços de comunicação e posicionamentos políticos. No entanto, ainda há lacunas na identificação de padrões psicolinguísticos em mandatos presidenciais brasileiros e na detecção de discursos destoantes. Para enfrentar essa lacuna, este trabalho propõe o SPEAR, um método que combina extração de dimensões do LIWC, remoção de outliers via LOF e agrupamento com K-Means, seguido da reintrodução dos discursos mais destoantes. Os resultados indicam melhora consistente da qualidade dos agrupamentos em 11 mandatos analisados.
Referências
Ahmadian, S., Azarshahi, S., and Paulhus, D. L. (2017). Explaining Donald Trump via communication style: Grandiosity, informality, and dynamism. Personality and Individual Differences, 107:49 – 53.
Bernard, N., Sagawa, Y., Bier, N., Lihoreau, T., Pazart, L. H., and Tannou, T. (2025). Using artificial intelligence for systematic review: the example of elicit. BMC Medical Research Methodology, 25.
Boyd, R. L. (2017). Psychological Text Analysis in the Digital Humanities. In Hai-Jew, S., editor, Data Analytics in Digital Humanities, pages 161–189. Springer International Publishing, Cham.
Breuniq, M. M., Kriegel, H.-P., Ng, R. T., and Sander, J. (2000). LOF: Identifying density-based local outliers. SIGMOD Record (ACM Special Interest Group on Management of Data), 29:93 – 104.
Carvalho, F., Junior, F. P., Ogasawara, E., Ferrari, L., and Guedes, G. (2024). Evaluation of the Brazilian Portuguese version of linguistic inquiry and word count 2015 (BP-LIWC2015). Language Resources and Evaluation, 58:203 – 222.
Carvalho, F., Rodrigues, R., Santos, G., Cruz, P., Ferrari, L., and Guedes, G. (2019). Avaliação da versão em português do LIWC Lexicon 2015 com análise de sentimentos em redes sociais. In Anais do VIII Brazilian Workshop on Social Network Analysis and Mining, pages 24–34, Porto Alegre, RS, Brasil. SBC.
Cezar, R. F. (2022). Brazilian Presidential Speeches from 1985 to July 2020.
de Brito Gadelha, S. R. (2011). Countercyclical fiscal policy, international financial crisis and economic growth in Brazil. Revista de Economia Politica, 31:794 – 812.
Gadelha, T., Monteiro, J. M., Machado, J., Claudino, I., Santos, R., Galick, L., and Santos, C. (2023). Ativismo da extrema direita brasileira no WhatsApp: O que mudou das eleições de 2018 para 2022? In Anais do XXXVIII Simpósio Brasileiro de Bancos de Dados, pages 336–341, Porto Alegre, RS, Brasil. SBC.
Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure.
Han, J., Kamber, M., and Pei, J. (2011). Data Mining: Concepts and Techniques. Morgan Kaufmann Publishers Inc., San Francisco, CA, United States, 3 edition.
Jain, A. K. (2010). Data clustering: 50 years beyond K-means. Pattern Recognition Letters, 31:651 – 666.
Jordan, K. N., Sterling, J., Pennebaker, J. W., and Boyd, R. L. (2019). Examining long-term trends in politics and culture through language of political leaders and cultural institutions. Proceedings of the National Academy of Sciences of the United States of America, 116:3476 – 3481.
Kangas, S. E. (2014). What can software tell us about political candidates? A critical analysis of a computerized method for political discourse. Journal of Language and Politics, 13:77 – 97.
Körner, R., Overbeck, J. R., Körner, E., and Schütz, A. (2022). How the Linguistic Styles of Donald Trump and Joe Biden Reflect Different Forms of Power. Journal of Language and Social Psychology, 41:631 – 658.
Lucas, T. P. B., Augusto, P., Reis, S. d., and Rocha, S. C. (2015). Impactos hidrometeóricos em Belo Horizonte-MG. Revista Brasileira de Climatologia, 16.
MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics, volume 5, pages 281–298. University of California press.
Malzer, C. and Baum, M. (2020). A Hybrid Approach To Hierarchical Density-based Cluster Selection. In 2020 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI), pages 223–228. IEEE.
McInnes, L., Healy, J., and Melville, J. (2020). UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.
Mimino, M. and Nakatani, S. (2021). langdetect: Port of Google’s language-detection library to Python.
Morgan, R. L., Whaley, P., Thayer, K. A., and Schünemann, H. J. (2018). Identifying the PECO: A framework for formulating good questions to explore the association of environmental and other exposures with health outcomes. Environment International, 121:1027 – 1031.
Nascimento, F., Monteiro, J. M., and Machado, J. (2025). Dinâmicas de Grupos de WhatsApp da Extrema Direita no Brasil: Uma Análise Comparativa Pré e Pós-Eleição de 2022. In Anais do XL Simpósio Brasileiro de Bancos de Dados, pages 563–575, Porto Alegre, RS, Brasil. SBC.
Okuno, H. Y., Carvalho, F., Guedes, G. P., and Torres, M. (2018). At analysis - Toward a psycholinguistic method to analyze video textual information. In CEUR Workshop Proceedings, volume 2170, pages 73 – 79.
Pennebaker, J. W., Boyd, R. L., Jordan, K., and Blackburn, K. (2015). The development and psychometric properties of LIWC2015. Technical report, University of Texas at Austin, Austin, TX.
Randour, F., Perrez, J., and Reuchamps, M. (2020). Twenty years of research on political discourse: A systematic review and directions for future research. Discourse and Society, 31:428 – 443.
Satopää, V., Albrecht, J., Irwin, D., and Raghavan, B. (2011). Finding a "kneedle"in a haystack: Detecting knee points in system behavior. In Proceedings - International Conference on Distributed Computing Systems, pages 166 – 171.
Shahapure, K. R. and Nicholas, C. (2020). Cluster quality analysis using silhouette score. In G, W., Z, Z., V.S, T., G, W., M, V., and L, C., editors, Proceedings - 2020 IEEE 7th International Conference on Data Science and Advanced Analytics, DSAA 2020, pages 747 – 748. Institute of Electrical and Electronics Engineers Inc.
Silva-Muller, L. and Sposito, H. (2024). Which Amazon Problem? Problem-constructions and Transnationalism in Brazilian Presidential Discourse since 1985. Environmental Politics, 33:398 – 421.
Sim, Y., Acree, B. D. L., Gross, J. H., and Smith, N. A. (2013). Measuring Ideological Proportions in Political Speeches. In Yarowsky, D., Baldwin, T., Korhonen, A., Livescu, K., and Bethard, S., editors, Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 91–101, Seattle, Washington, USA. Association for Computational Linguistics.
Sposito, H. (2025). Radiating Truthiness: Authenticity Performances in Politics in Brazil and the United States. Political Studies, 73:770 – 794.
Stromer-Galley, J., Rossini, P., Hemsley, J., Bolden, S. E., and McKernan, B. (2021). Political Messaging Over Time: A Comparison of US Presidential Candidate Facebook Posts and Tweets in 2016 and 2020. Social Media and Society, 7.
Tausczik, Y. R. and Pennebaker, J. W. (2010). The psychological meaning of words: LIWC and computerized text analysis methods. Journal of Language and Social Psychology, 29:24 – 54.
Vilela, E. and Neiva, P. (2011). Temas e regiões nas políticas externas de Lula e Fernando Henrique: comparação do discurso dos dois presidentes. Revista Brasileira de Política Internacional, 54:70–96.
Wilkinson, L. (2006). Statistical computing and graphics: Revising the Pareto chart. American Statistician, 60:332 – 334.
