Which Feature-Classifier Combinations for Video Sign Language Recognition?
Resumo
This paper presents a comparative investigation of five feature extractors frequently used for video-based sign language recognition: I3D-RGB, I3D-OF, MediaPipe, Pose, TSM, and SlowFast. Many studies utilize these features, yet there is a lack of comparative analysis regarding their complexity and discriminative power when combined with linear, non-linear, and deep neural classifiers. Test results based on the public UFOP-LIBRAS-ISO dataset indicate that the top two performances were achieved by the TMS+LR combination (F1-score of 99.30%) and the SlowFast+LR combination (F1-score of 98.40%). Conversely, the complexity ranking begins with Pose as the least computationally demanding, followed by TSM, I3D-RGB, I3D-OF, MediaPipe, and SlowFast, with the latter being the most complex.
Referências
Boháček, M. and Hrúz, M. Sign pose-based transformer for word-level sign language recognition. In Proceedings of the IEEE/CVF winter conference on applications of computer vision. pp. 182–191, 2022.
Carreira, J. and Zisserman, A. Quo vadis, action recognition? a new model and the kinetics dataset. [link], 2018.
Cerna, L. R., Cardenas, E. E., Miranda, D. G., Menotti, D., and Camara-Chavez, G. A multimodal libras-ufop brazilian sign language dataset of minimal pairs using a microsoft kinect sensor. Expert Systems with Applications vol. 167, pp. 114179, 2021.
Desai, A., Berger, L., Minakov, F. O., Milan, V., Singh, C., Pumphrey, K., Ladner, R. E., Daumé III, H., Lu, A. X., Caselli, N., and Bragg, D. ASL citizen: A community-sourced dataset for advancing isolated sign language recognition. In arXiv preprint, 2023.
Fawcett, T. An introduction to ROC analysis. Pattern Recognition Letters 27 (8): 861–874, 2006.
Feichtenhofer, C., Fan, H., Malik, J., and He, K. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 6202–6211, 2019.
Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation 9 (8): 1735–1780, 1997.
Kuznetsova, A. and Kimmelman, V. Testing MediaPipe holistic for linguistic analysis of nonmanual markers in sign languages, 2024.
LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. Nature 521 (7553): 436–444, 2015.
Li, D., Rodriguez Opazo, C., Yu, X., and Li, H. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. arXiv preprint arXiv:1910.11006, 2020.
Lin, J., Gan, C., and Han, S. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE International Conference on Computer Vision, 2019.
Lugaresi, C., Tang, J., Nash, H., McClanahan, C., Uboweja, E., Hays, M., Zhang, F., Chang, C.-L., Yong, M. G., Lee, J., Chang, W.-T., Hua, W., Georg, M., and Grundmann, M. MediaPipe: A framework for building perception pipelines, 2019.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems (NeurIPS). Vol. 32. Curran Associates, Inc., pp. 8024–8035, 2019.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, Scikit-learn: Machine learning in Python. Journal of Machine Learning Research vol. 12, pp. 2825–2830, 2011.
Renjith, S., Varghese, A., Rashmi, M., and Poorna, S. Transformer-based motion-visual integrated fusion for isolated sign language recognition. Computers and Electrical Engineering vol. 130, pp. 110902, 2026.
Sarhan, N. and Frintrop, S. Transfer learning for videos: From action recognition to sign language recognition. In 2020 IEEE International Conference on Image Processing (ICIP). IEEE, pp. 1811–1815, 2020.
Singh, S., Keserwani, P., Inoue, K., Iwamura, M., and Roy, P. P. Revisiting i3d for sign language recognition. IEICE Transactions on Information and Systems, 2025.
van Rijsbergen, C. J. Information retrieval. Butterworth, 1979.
Zhu, Q., Li, J., Yuan, F., and Gan, Q. Temporal superimposed crossover module for effective continuous sign language. arXiv preprint arXiv:2211.03387, 2022.
