Which Feature-Classifier Combinations for Video Sign Language Recognition?

  • Luan Fernandes De Franca UECE
  • José Everardo Bessa Maia UECE

Resumo


This paper presents a comparative investigation of five feature extractors frequently used for video-based sign language recognition: I3D-RGB, I3D-OF, MediaPipe, Pose, TSM, and SlowFast. Many studies utilize these features, yet there is a lack of comparative analysis regarding their complexity and discriminative power when combined with linear, non-linear, and deep neural classifiers. Test results based on the public UFOP-LIBRAS-ISO dataset indicate that the top two performances were achieved by the TMS+LR combination (F1-score of 99.30%) and the SlowFast+LR combination (F1-score of 98.40%). Conversely, the complexity ranking begins with Pose as the least computationally demanding, followed by TSM, I3D-RGB, I3D-OF, MediaPipe, and SlowFast, with the latter being the most complex.

Palavras-chave: LIBRAS, Sign Language Recognition, Transfer Learning, Deep Learning, Feature Extraction

Referências

Ahn, J., Jang, Y., and Chung, J. S. Slowfast network for continuous sign language recognition. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 3920–3924, 2024.

Boháček, M. and Hrúz, M. Sign pose-based transformer for word-level sign language recognition. In Proceedings of the IEEE/CVF winter conference on applications of computer vision. pp. 182–191, 2022.

Carreira, J. and Zisserman, A. Quo vadis, action recognition? a new model and the kinetics dataset. [link], 2018.

Cerna, L. R., Cardenas, E. E., Miranda, D. G., Menotti, D., and Camara-Chavez, G. A multimodal libras-ufop brazilian sign language dataset of minimal pairs using a microsoft kinect sensor. Expert Systems with Applications vol. 167, pp. 114179, 2021.

Desai, A., Berger, L., Minakov, F. O., Milan, V., Singh, C., Pumphrey, K., Ladner, R. E., Daumé III, H., Lu, A. X., Caselli, N., and Bragg, D. ASL citizen: A community-sourced dataset for advancing isolated sign language recognition. In arXiv preprint, 2023.

Fawcett, T. An introduction to ROC analysis. Pattern Recognition Letters 27 (8): 861–874, 2006.

Feichtenhofer, C., Fan, H., Malik, J., and He, K. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 6202–6211, 2019.

Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation 9 (8): 1735–1780, 1997.

Kuznetsova, A. and Kimmelman, V. Testing MediaPipe holistic for linguistic analysis of nonmanual markers in sign languages, 2024.

LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. Nature 521 (7553): 436–444, 2015.

Li, D., Rodriguez Opazo, C., Yu, X., and Li, H. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. arXiv preprint arXiv:1910.11006, 2020.

Lin, J., Gan, C., and Han, S. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE International Conference on Computer Vision, 2019.

Lugaresi, C., Tang, J., Nash, H., McClanahan, C., Uboweja, E., Hays, M., Zhang, F., Chang, C.-L., Yong, M. G., Lee, J., Chang, W.-T., Hua, W., Georg, M., and Grundmann, M. MediaPipe: A framework for building perception pipelines, 2019.

Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems (NeurIPS). Vol. 32. Curran Associates, Inc., pp. 8024–8035, 2019.

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, Scikit-learn: Machine learning in Python. Journal of Machine Learning Research vol. 12, pp. 2825–2830, 2011.

Renjith, S., Varghese, A., Rashmi, M., and Poorna, S. Transformer-based motion-visual integrated fusion for isolated sign language recognition. Computers and Electrical Engineering vol. 130, pp. 110902, 2026.

Sarhan, N. and Frintrop, S. Transfer learning for videos: From action recognition to sign language recognition. In 2020 IEEE International Conference on Image Processing (ICIP). IEEE, pp. 1811–1815, 2020.

Singh, S., Keserwani, P., Inoue, K., Iwamura, M., and Roy, P. P. Revisiting i3d for sign language recognition. IEICE Transactions on Information and Systems, 2025.

van Rijsbergen, C. J. Information retrieval. Butterworth, 1979.

Zhu, Q., Li, J., Yuan, F., and Gan, Q. Temporal superimposed crossover module for effective continuous sign language. arXiv preprint arXiv:2211.03387, 2022.
Publicado
19/10/2026
FRANCA, Luan Fernandes De; MAIA, José Everardo Bessa. Which Feature-Classifier Combinations for Video Sign Language Recognition?. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 273-280. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.31025.