Classical Tricks for Better Topics: Can Traditional Text Mining Techniques Improve Generative Topic Modeling?
Resumo
Topic modeling is widely used to discover and organize latent themes in large text collections. While recent embedding-based methods have improved topic discovery by leveraging dense semantic representations, they typically represent topics as lists of keywords, often providing limited insights into the underlying theme. More recently, Large Language Model (LLM)-based approaches, such as TopicGPT, have introduced generative topic modeling methods capable of producing human-readable topic labels and descriptions. However, the effectiveness of these methods depends heavily on the quality and representativeness of the documents selected to guide topic generation. In this paper, we propose Core-Periphery Sampling (CPS), a hybrid strategy that combines classical text representation and clustering methods with a TopicGPT-inspired pipeline. CPS selects both central documents and boundary instances that capture transitional and potentially informative cases, providing the generative model with an informed and diverse subset for topic induction. Experimental results show that CPS outperforms random sampling, particularly when the underlying document representation yields well-separable clusters. Although BERTopic remains stronger on clustering-oriented metrics, CPS offers an effective, computationally simple mechanism for guiding LLM-based topic generation via informed document selection, resulting in coherent and contextually grounded topic descriptions.
Referências
Alibaba Cloud. Qwen3 technical report, 2025.
Amorim, A., Murrugarra-Llerena, N., Silva, V., de Oliveira, D., and Paes, A. ALTES: uma ferramenta de rotulação automática de tópicos por meio de fontes externas. In Anais Estendidos do XXXVIII Simpósio Brasileiro de Bancos de Dados. SBC, Porto Alegre, RS, Brasil, pp. 120–125, 2023.
Angelov, D. Top2vec: Distributed representations of topics, 2020.
Anh-Hoang, D., Tran, V., and Nguyen, L.-M. Survey and analysis of hallucinations in large language models: attribution to prompting strategies or model behavior. Frontiers in Artificial Intelligence vol. 8, pp. 1622292, 2025.
Blei, D. M. Probabilistic topic models. Commun. ACM 55 (4): 77–84, Apr., 2012.
Blei, D. M., Ng, A. Y., and Jordan, M. I. Latent dirichlet allocation. J. Mach. Learn. Res. vol. 3, Mar., 2003.
Campello, R., Moulavi, D., Zimek, A., and Sander, J. Hierarchical density estimates for data clustering, visualization, and outlier detection. ACM Transactions on Knowledge Discovery from Data 10 (1): 1–51, 2015.
Chuang, J., Gupta, S., Manning, C., and Heer, J. Topic model diagnostics: Assessing domain relevance via topical alignment. In Proceedings of the 30th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 28. PMLR, Atlanta, Georgia, USA, pp. 612–620, 2013.
Deerwester, S., Dumais, S. T., Furnas, G. W., Landauer, T. K., and Harshman, R. Indexing by latent semantic analysis. Journal of the American Society for Information Science 41 (6): 391–407, 1990.
Grootendorst, M. BERTopic: Neural topic modeling with a class-based tf-idf procedure, 2022.
Hoyle, A., Goel, P., Peskov, D., Hian-Cheong, A., Boyd-Graber, J., and Resnik, P. Is automated topic model evaluation broken? the incoherence of coherence. In Proceedings of the 35th International Conference on Neural Information Processing Systems. NIPS ’21. Curran Associates Inc., Red Hook, NY, USA, 2021.
Hubert, L. and Arabie, P. Comparing partitions. Journal of Classification 2 (1): 193–218, Dec, 1985.
Lee, D. and Seung, H. S. Algorithms for non-negative matrix factorization. In Advances in Neural Information Processing Systems. Vol. 13. MIT Press, 2000.
MacQueen, J. B. Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability. Vol. 1. UC Press, Berkeley, CA, 1967.
Manning, C., Raghavan, P., and Schütze, H. Introduction to Information Retrieval. Cambridge UP, 2008.
Meng, Y., Zhang, Y., Huang, J., Zhang, Y., and Han, J. Topic discovery via latent space clustering of pretrained language model representations. In Proceedings of the ACM web conference 2022. pp. 3143–3152, 2022.
Merity, S., Keskar, N. S., and Socher, R. Regularizing and optimizing LSTM language models. In International Conference on Learning Representations, 2018.
Mu, Y., Dong, C., Bontcheva, K., and Song, X. Large language models offer an alternative to the traditional approach of topic modelling. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA and ICCL, Torino, Italia, pp. 10160–10171, 2024.
Pham, C. M., Hoyle, A., Sun, S., Resnik, P., and Iyyer, M. TopicGPT: A prompt-based topic modeling framework. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, Mexico City, Mexico, pp. 2956–2984, 2024.
Röder, M., Both, A., and Hinneburg, A. Exploring the space of topic coherence measures. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining. WSDM ’15. Association for Computing Machinery, New York, NY, USA, pp. 399–408, 2015.
Rousseeuw, P. J. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics vol. 20, pp. 53–65, 1987.
Strehl, A. and Ghosh, J. Cluster ensembles — a knowledge reuse framework for combining multiple partitions. J. Mach. Learn. Res. 3 (null): 583–617, Mar., 2003.
van der Maaten, L. and Hinton, G. Visualizing data using t-SNE. Journal of Machine Learning Research, 2008.
Vinh, N. X., Epps, J., and Bailey, J. Information theoretic measures for clusterings comparison: Variants, properties, normalization and correction for chance. Journal of Machine Learning Research 11 (95): 2837–2854, 2010.
Wang, H., Prakash, N., Hoang, N. K., Hee, M. S., Naseem, U., and Lee, R. K.-W. Prompting large language models for topic modeling. In 2023 IEEE International Conference on Big Data (BigData). pp. 1236–1241, 2023.
Yang, X., Zhao, H., Xu, W., Qi, Y., Lu, J., Phung, D., and Du, L. Neural topic modeling with large language models in the loop. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Vienna, Austria, pp. 1377–1401, 2025.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. BERTScore: Evaluating text generation with bert. In International Conference on Learning Representations, 2020.
Zhao, Y. and Karypis, G. Criterion functions for document clustering, 2001.
