Knowledge Discovery of Emerging Developer Skill Profiles in Telegram Community Using Semantic Embeddings
Resumo
This article proposes a pipeline to discover professional profiles of software developers from job posts published on Telegram. The corpus consists of 22,243 messages collected between May 2021 and January 2026 from a public Brazilian group dedicated to IT vacancies; a two-stage filter combining deterministic rules and semantic validation reduced the set to 7,264 valid posts. A taxonomy of 108 technical and contractual competences was applied to the corpus, three embedding models were compared, and the reduced space was partitioned with K-Means. The number of clusters followed internal indices restricted by admissibility criteria defined beforehand and decided by an independent evaluator, resulting in 23 profiles with mean topic coherence of 0.522. Labelling and auditing relied on models from different families, so that no model judged its own output. This choice mattered: the cross-architecture protocol spread the scores instead of concentrating them, and showed that about a fifth of the corpus forms clusters built around posting format rather than occupational content. With each competence expressed as a share of monthly posts and multiple-testing correction applied, only three of 50 series decline significantly and none rises.
Referências
Batista, J. V. C. d. R., de Meireles Costa, G., and de Freitas Jorge, E. M. Mecanismo de busca semântica baseado em word embeddings em dados do currículo lattes, programas de pós-graduação e grupos de pesquisa. In Escola Regional de Computação Bahia, Alagoas e Sergipe (ERBASE). SBC, pp. 109–118, 2024.
Benjamini, Y. and Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal statistical society: series B (Methodological) 57 (1): 289–300, 1995.
Boselli, R., Cesarini, M., Mercorio, F., and Mezzanzanica, M. Using machine learning for labour market intelligence. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, pp. 330–342, 2017.
Colace, F., Santo, M. D., Lombardi, M., Mercorio, F., Mezzanzanica, M., and Pascale, F. Towards labour market intelligence through topic modelling. In Proceedings of the 52nd Hawaii International Conference on System Sciences. pp. 5256–5265, 2019.
Debao, D., Yinxia, M., and Min, Z. Analysis of big data job requirements based on k-means text clustering in china. PloS one 16 (8): e0255419, 2021.
Doğan, Y., Dalkılıç, F., Kut, A., Kara, K. C., and Takazoğlu, U. A novel stream mining approach as streamcluster feature tree algorithm: A case study in turkish job postings. Applied Sciences 12 (15): 7893, 2022.
Grootendorst, M. Bertopic: Neural topic modeling with a class-based tf-idf procedure. arXiv preprint arXiv:2203.05794 , 2022.
Ozcan, S., Sakar, C. O., and Suloglu, M. Human resources mining for examination of r&d progress and requirements. IEEE Transactions on Engineering Management 68 (5): 1372–1387, 2020.
Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). pp. 3982–3992, 2019.
Sibarani, E. M. and Scerri, S. Scodis: Job advert-derived time series for high-demand skillset discovery and prediction. In International Conference on Database and Expert Systems Applications. Springer, pp. 366–381, 2020.
Siswipraptini, P. C., Warnars, H. L. H. S., Ramadhan, A., and Budiharto, W. Information technology job profile using average-linkage hierarchical clustering analysis. IEEE Access vol. 11, pp. 94647–94663, 2023.
Souza, F., Nogueira, R., and Lotufo, R. Bertimbau: pretrained bert models for brazilian portuguese. In Brazilian conference on intelligent systems. Springer, pp. 403–417, 2020.
Sun, L., Cao, X., and Su, D. Statistical analysis of online recruitment information based on text mining technology. In 2022 International Conference on Intelligent Transportation, Big Data & Smart City (ICITBS). IEEE, pp. 431–436, 2022.
Wibowo, Y. P. Exploring job vacancy topics and trends in indonesia using latent dirichlet allocation (lda) and exploratory data analysis (eda). Journal of Artificial Intelligence and Engineering Applications (JAIEA) 5 (1): 1778–1785, 2025.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems vol. 36, pp. 46595–46623, 2023.
