Unsupervised behavioral profiling of deception (bluffing) using structured and embedding features with clustering algorithm
Resumo
Introduction: This work investigates the use of Machine Learning techniques to analyze bluffing behavior in the social deduction game Ultimate Werewolf. Objective: To identify and characterize behavioral profiles of bluffing players, addressing challenges such as temporality, mixed data types, and the non-exclusive nature of behavioral patterns. Methodology or Steps: A pipeline that combines data preprocessing, Large Language Model for bluff classification, feature engineering with temporality and interaction attributes, and clustering. Multiple experimental scenarios were evaluated, including structured features, embedding-based representations, and hybrid approaches. Results: Although not achieving higher internal validation scores, experiments show that combining feature filtering with embedding-based representations and Gower distance leads to more coherent and interpretable clusters. These findings indicate that bluffing behavior is not homogeneous, but instead organized into identifiable patterns, highlighting the effectiveness of the proposed approach for analyzing behavior in conversation-driven games.
Palavras-chave:
Machine Learning, Social Deduction Games, Role-Playing Games (RPG), Bluff, Player Behavior Modeling
Referências
Bauckhage, C., Drachen, A., e Sifa, R. (2015). Clustering game behavior data. IEEE Transactions on Computational Intelligence and AI in Games, 7(3):266–278.
Ben Ali, B. e Massmoudi, Y. (2013). K-means clustering based on gower similarity coefficient: A comparative study. In 2013 5th International Conference on Modeling, Simulation and Applied Optimization (ICMSAO), pages 1–5.
Cachay-Guivin, A. (2024). Evaluation on embeddings application for spanish automatic text clustering. Ingeniare. Revista chilena de ingeniería, 32.
Damle, A., Minden, V., e Ying, L. (2019). Simple, direct and efficient multi-way spectral clustering. Information and Inference: A Journal of the IMA, 8(1):181–203.
Davies, D. L. e Bouldin, D. W. (1979). A cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-1(2):224–227.
de Amorim, L. B., Cavalcanti, G. D., e Cruz, R. M. (2023). The choice of scaling technique matters for classification performance. Applied Soft Computing, 133:109924.
Ding, L., Li, C., Jin, D., e Ding, S. (2024). Survey of spectral clustering based on graph theory. Pattern Recognition, 151:110366.
Drachen, A., Sifa, R., Bauckhage, C., e Thurau, C. (2012). Guns, swords and data: Clustering of player behavior in computer games in the wild. In 2012 IEEE Conference on Computational Intelligence and Games (CIG), pages 163–170.
Drachen, A., Thurau, C., Sifa, R., e Bauckhage, C. (2014). A comparison of methods for player clustering via behavioral telemetry. CoRR, abs/1407.3950.
Field, A., Miles, J., e Field, Z. (2012). Discovering Statistics Using R. SAGE Publications, 1 edition.
Games, B. (2024). Ultimate werewolf.
Gower, J. C. (1971). A general coefficient of similarity and some of its properties. Biometrics, 27(4):857–871.
Grande-de Prado, M., Baelo, R., García-Martín, S., e Abella-García, V. (2020). Mapping role-playing games in ibero-america: An educational review. Sustainability, 12(16).
Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., Wang, S., Zhang, K., Lin, Z., Zhang, B., Ni, L. M., Gao, W., Wang, Y., e Guo, J. (2026). A survey on llm-as-a-judge. The Innovation, 7(6):101253.
Hagendorff, T. (2024). Deception abilities emerged in large language models. Proceedings of the National Academy of Sciences, 121(24):e2317967121.
Harley, J. B. e Sparkman, D. (2019). Machine learning and nde: Past, present, and future. AIP Conference Proceedings, 2102(1):090001.
Hope, T. M. (2020). Chapter 4 - linear regression. In Mechelli, A. e Vieira, S., editors, Machine Learning, pages 67–81. Academic Press.
Hu, C., Zhao, Y., Wang, Z., Du, H., e Liu, J. (2024). Games for artificial intelligence research: A review and perspectives. IEEE Transactions on Artificial Intelligence, 5(12):5949–5968.
Hubert, L. e Arabie, P. (1985). Comparing partitions. Journal of Classification, 2(1):193–218.
Jaques, R. R. e Bisol, C. A. (2019). Role-playing games for common well-being in high school — notes from a study developed in southern brazil. International Journal of Information and Education Technology, 9(7):498–501.
Jolliffe, I. T. (2002). Principal Component Analysis. Springer, 2 edition.
Lai, B., Zhang, H., Liu, M., Pariani, A., Ryan, F., Jia, W., Hayati, S. A., Rehg, J., e Yang, D. (2023). Werewolf among us: Multimodal resources for modeling persuasion behaviors in social deduction games. In Rogers, A., Boyd-Graber, J., e Okazaki, N., editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 6570–6588, Toronto, Canada. Association for Computational Linguistics.
LeCun, Y., Bengio, Y., e Hinton, G. E. (2015). Deep learning. Nature, 521(7553):436–444.
Malzer, C. e Baum, M. (2019). Hdbscan(ϵ ): An alternative cluster extraction method for HDBSCAN. CoRR, abs/1911.02282.
Manning, C. D., Raghavan, P., e Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.
Marwala, T. e Hurwitz, E. (2009). A multi-agent approach to bluffing. In Multiagent systems. Citeseer.
Maulud, D. H. e Abdulazeez, A. M. (2020). A review on linear regression comprehensive in machine learning. Journal of Applied Science and Technology Trends, 1(2):140–147.
Mottaqi, M. S., Mohammadipanah, F., e Sajedi, H. (2021). Contribution of machine learning approaches in response to sars-cov-2 infection. Informatics in Medicine Unlocked, 23:100526.
Murphy, A. H. (1996). The finley affair: A signal event in the history of forecast verification. Weather and Forecasting, 11(1):3 – 20.
Perillo, J. T. e Kassin, S. M. (2011). Inside interrogation: The lie, the bluff, and false confessions. Law and Human Behavior, 35(4):327–337.
Rath, S., Tripathy, A., e Tripathy, A. R. (2020). Prediction of new active cases of coronavirus disease (covid-19) pandemic using multiple linear regression model. Diabetes Metabolic Syndrome: Clinical Research Reviews, 14(5):1467–1474.
Sun, L. (2025). slhleosun/werewolfgameplays.
Tan, P.-N., Steinbach, M., Karpatne, A., e Kumar, V. (2018). Introduction to Data Mining (2nd Edition). Pearson, 2nd edition.
van der Maaten, L. e Hinton, G. (2008). Visualizing data using t-sne. Journal of Machine Learning Research, 9:2579–2605.
Wang, L., Yang, N., Huang, X., Yang, L., Majumder, R., e Wei, F. (2024). Multilingual e5 text embeddings: A technical report.
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., e Zhou, M. (2020). Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.
Wang, Y.-H., Zhang, Y.-F., Zhang, Y., Gu, Z.-F., Zhang, Z.-Y., Lin, H., e Deng, K.-J. (2022). Identification of adaptor proteins using the anova feature selection technique. Methods, 208:42–47.
Yoo, B. e Kim, K.-J. (2024). Finding deceivers in social context with large language models and how to find them: the case of the mafia game. Scientific Reports, 14(1):30946.
Øyvind Langsrud (2003). Anova for unbalanced data: Use type ii instead of type iii sums of squares. Statistics and Computing, 13(2):163–167.
Ben Ali, B. e Massmoudi, Y. (2013). K-means clustering based on gower similarity coefficient: A comparative study. In 2013 5th International Conference on Modeling, Simulation and Applied Optimization (ICMSAO), pages 1–5.
Cachay-Guivin, A. (2024). Evaluation on embeddings application for spanish automatic text clustering. Ingeniare. Revista chilena de ingeniería, 32.
Damle, A., Minden, V., e Ying, L. (2019). Simple, direct and efficient multi-way spectral clustering. Information and Inference: A Journal of the IMA, 8(1):181–203.
Davies, D. L. e Bouldin, D. W. (1979). A cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-1(2):224–227.
de Amorim, L. B., Cavalcanti, G. D., e Cruz, R. M. (2023). The choice of scaling technique matters for classification performance. Applied Soft Computing, 133:109924.
Ding, L., Li, C., Jin, D., e Ding, S. (2024). Survey of spectral clustering based on graph theory. Pattern Recognition, 151:110366.
Drachen, A., Sifa, R., Bauckhage, C., e Thurau, C. (2012). Guns, swords and data: Clustering of player behavior in computer games in the wild. In 2012 IEEE Conference on Computational Intelligence and Games (CIG), pages 163–170.
Drachen, A., Thurau, C., Sifa, R., e Bauckhage, C. (2014). A comparison of methods for player clustering via behavioral telemetry. CoRR, abs/1407.3950.
Field, A., Miles, J., e Field, Z. (2012). Discovering Statistics Using R. SAGE Publications, 1 edition.
Games, B. (2024). Ultimate werewolf.
Gower, J. C. (1971). A general coefficient of similarity and some of its properties. Biometrics, 27(4):857–871.
Grande-de Prado, M., Baelo, R., García-Martín, S., e Abella-García, V. (2020). Mapping role-playing games in ibero-america: An educational review. Sustainability, 12(16).
Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., Wang, S., Zhang, K., Lin, Z., Zhang, B., Ni, L. M., Gao, W., Wang, Y., e Guo, J. (2026). A survey on llm-as-a-judge. The Innovation, 7(6):101253.
Hagendorff, T. (2024). Deception abilities emerged in large language models. Proceedings of the National Academy of Sciences, 121(24):e2317967121.
Harley, J. B. e Sparkman, D. (2019). Machine learning and nde: Past, present, and future. AIP Conference Proceedings, 2102(1):090001.
Hope, T. M. (2020). Chapter 4 - linear regression. In Mechelli, A. e Vieira, S., editors, Machine Learning, pages 67–81. Academic Press.
Hu, C., Zhao, Y., Wang, Z., Du, H., e Liu, J. (2024). Games for artificial intelligence research: A review and perspectives. IEEE Transactions on Artificial Intelligence, 5(12):5949–5968.
Hubert, L. e Arabie, P. (1985). Comparing partitions. Journal of Classification, 2(1):193–218.
Jaques, R. R. e Bisol, C. A. (2019). Role-playing games for common well-being in high school — notes from a study developed in southern brazil. International Journal of Information and Education Technology, 9(7):498–501.
Jolliffe, I. T. (2002). Principal Component Analysis. Springer, 2 edition.
Lai, B., Zhang, H., Liu, M., Pariani, A., Ryan, F., Jia, W., Hayati, S. A., Rehg, J., e Yang, D. (2023). Werewolf among us: Multimodal resources for modeling persuasion behaviors in social deduction games. In Rogers, A., Boyd-Graber, J., e Okazaki, N., editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 6570–6588, Toronto, Canada. Association for Computational Linguistics.
LeCun, Y., Bengio, Y., e Hinton, G. E. (2015). Deep learning. Nature, 521(7553):436–444.
Malzer, C. e Baum, M. (2019). Hdbscan(ϵ ): An alternative cluster extraction method for HDBSCAN. CoRR, abs/1911.02282.
Manning, C. D., Raghavan, P., e Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.
Marwala, T. e Hurwitz, E. (2009). A multi-agent approach to bluffing. In Multiagent systems. Citeseer.
Maulud, D. H. e Abdulazeez, A. M. (2020). A review on linear regression comprehensive in machine learning. Journal of Applied Science and Technology Trends, 1(2):140–147.
Mottaqi, M. S., Mohammadipanah, F., e Sajedi, H. (2021). Contribution of machine learning approaches in response to sars-cov-2 infection. Informatics in Medicine Unlocked, 23:100526.
Murphy, A. H. (1996). The finley affair: A signal event in the history of forecast verification. Weather and Forecasting, 11(1):3 – 20.
Perillo, J. T. e Kassin, S. M. (2011). Inside interrogation: The lie, the bluff, and false confessions. Law and Human Behavior, 35(4):327–337.
Rath, S., Tripathy, A., e Tripathy, A. R. (2020). Prediction of new active cases of coronavirus disease (covid-19) pandemic using multiple linear regression model. Diabetes Metabolic Syndrome: Clinical Research Reviews, 14(5):1467–1474.
Sun, L. (2025). slhleosun/werewolfgameplays.
Tan, P.-N., Steinbach, M., Karpatne, A., e Kumar, V. (2018). Introduction to Data Mining (2nd Edition). Pearson, 2nd edition.
van der Maaten, L. e Hinton, G. (2008). Visualizing data using t-sne. Journal of Machine Learning Research, 9:2579–2605.
Wang, L., Yang, N., Huang, X., Yang, L., Majumder, R., e Wei, F. (2024). Multilingual e5 text embeddings: A technical report.
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., e Zhou, M. (2020). Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.
Wang, Y.-H., Zhang, Y.-F., Zhang, Y., Gu, Z.-F., Zhang, Z.-Y., Lin, H., e Deng, K.-J. (2022). Identification of adaptor proteins using the anova feature selection technique. Methods, 208:42–47.
Yoo, B. e Kim, K.-J. (2024). Finding deceivers in social context with large language models and how to find them: the case of the mafia game. Scientific Reports, 14(1):30946.
Øyvind Langsrud (2003). Anova for unbalanced data: Use type ii instead of type iii sums of squares. Statistics and Computing, 13(2):163–167.
Publicado
29/09/2026
Como Citar
BERTOLDO, Kevin S.; JULIA, Rita M. S.; FARIA, Elaine R.; NASCIMENTO, Marcelo Z. do.
Unsupervised behavioral profiling of deception (bluffing) using structured and embedding features with clustering algorithm. In: SIMPÓSIO BRASILEIRO DE JOGOS E ENTRETENIMENTO DIGITAL (SBGAMES), 25. , 2026, Goiânia/GO.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 991-1003.
DOI: https://doi.org/10.5753/sbgames.2026.26222.
