LLM-based Description Enrichment for Short Video Clustering
Resumo
The rapid growth of short-video platforms has created large collections of visual content that are difficult to organize manually. Video clustering offers a practical way to explore these collections, but its quality depends on representations that capture the meaning of each video, not only its visual appearance. Compact vision-language models can generate useful descriptions, although these descriptions are often too generic to support fine-grained semantic grouping. We introduce a method for improving unsupervised video clustering by building richer semantic representations from short visual descriptions before embedding generation. Instead of relying directly on the initial captions produced by a compact vision-language model, the method expands each description with information about actions, objects, context, intent, scene type, and possible high-level categories. Multiple Large Language Models are used to generate complementary semantic views of the same video, which are then converted into embeddings and clustered without using labels. Experiments on the MSR-VTT benchmark show that enriched descriptions produce more coherent clusters than the original descriptions: the best overall configuration (Claude-enriched descriptions with OpenAI embeddings) reaches an NMI of 0.3740, a 5.71% relative gain over its caption-only baseline, while the largest relative improvement (13.18%) is obtained with Qwen embeddings. The results indicate that LLM-based enrichment helps reduce the semantic gap between visual content and high-level meaning, offering a practical path to organize unlabeled video collections under a known taxonomy.
Referências
Ding, K., Wang, Y., Liu, P., Yu, Q., Zhang, H., Xiang, S., and Pan, C. Prompt tuning with soft context sharing for vision-language models, 2024.
Efron, B. and Tibshirani, R. J. An Introduction to the Bootstrap. Chapman and Hall/CRC, London, UK, 1993.
Fan, L., Krishnan, D., Isola, P., Katabi, D., and Tian, Y. Improving clip training with language rewrites. In Advances in Neural Information Processing Systems. Vol. 36. New Orleans, USA, 2023.
Huang, W., Wu, A., Yang, Y., Luo, X., Yang, Y., Hu, L., Dai, Q., Dai, X., Chen, D., Luo, C., and Qiu, L. Llm2clip: Powerful language model unlocks richer cross-modality representation, 2024.
Hubert, L. and Arabie, P. Comparing partitions. Journal of Classification 2 (1): 193–218, 1985.
Hugging Face. Smolvlm2: A small vision language model. [link], 2024.
Jia, M., Tang, L., Cheng, B., Xu, S., Ding, S., and Dai, J. Visual prompt tuning. In European Conference on Computer Vision. Tel Aviv, Israel, 2022.
Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning, 2023.
Liu, P., Wang, X., Cui, Z., and Ye, W. Queries are not alone: Clustering text embeddings for video search, 2025.
Luo, H., Ji, L., Zhong, M., Chen, Y., Lei, W., Duan, N., and Li, T. Clip4clip: An empirical study of clip for end to end video clip retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal, Canada, pp. 11451–11461, 2021.
Menon, S. and Vondrick, C. Visual classification via description from large language models. In Proceedings of the Eleventh International Conference on Learning Representations (ICLR). Kigali, Rwanda, 2023.
Molem, A., Makri, S., and Mckay, D. Keepin’ it reel: Investigating how short videos on tiktok and instagram reels influence view change. In Proceedings of the 2024 ACM SIGIR Conference on Human Information Interaction and Retrieval. Association for Computing Machinery, New York, NY, USA, pp. 317–327, 2024.
OpenRouter. Openrouter: Unified api for language models. [link], 2024.
Pratt, S., Covert, I., Liu, R., and Farhadi, A. What does a platypus look like? generating customized prompts for zero-shot image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 15691–15701, 2023.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning. Virtual, pp. 8748–8763, 2021.
Shen, Y., Mao, Z., and Xu, H. Multitask vision-language prompt tuning, 2024.
Smeulders, A. W. M., Worring, M., Santini, S., Gupta, A., and Jain, R. Content-based image retrieval at the end of the early years. IEEE Transactions on Pattern Analysis and Machine Intelligence 22 (12): 1349–1380, 2000.
Strehl, A. and Ghosh, J. Cluster ensembles—a knowledge reuse framework for combining multiple partitions. Journal of Machine Learning Research vol. 3, pp. 583–617, 2002.
Van Daele, T., Iyer, A., Zhang, Y., Derry, J. C., Huh, M., and Pavel, A. Making short-form videos accessible with hierarchical video summaries. In Proceedings of the CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 2024.
Wang, Z., Mi, S., and Zhang, Y. Dvc2: Deep video cascade clustering from video structures. Neurocomputing vol. 657, 2025.
Wilcoxon, F. Individual comparisons by ranking methods. Biometrics Bulletin 1 (6): 80–83, 1945.
Xu, J., Mei, T., Yao, T., and Rui, Y. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5288–5296, 2016.
Zannettou, S., Nemes-Nemeth, O., Ayalon, O., Goetzen, A., Gummadi, K. P., Redmiles, E. M., and Roesner, F. Analyzing user engagement with tiktok’s short format video recommendations using data donations. In Proceedings of the CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 2024.
Zhang, H. et al. Unsupervised multimodal clustering for semantics discovery in multimodal utterances, 2024.
Zhang, J., Huang, J., Jin, S., and Lu, S. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (3): 1449–1469, 2024.
Zhou, K., Yang, J., Loy, C. C., and Liu, Z. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans, USA, pp. 16816–16825, 2022a.
Zhou, K., Yang, J., Loy, C. C., and Liu, Z. Learning to prompt for vision-language models. International Journal of Computer Vision vol. 130, pp. 2337–2348, 2022b.
