Call For Paper - Upcoming Conferences

Research Article | Open Access | Download PDF
Volume 13 | Issue 7 | Year 2026 | Article Id. IJECE-V13I7P109 | DOI : https://doi.org/10.14445/23488549/IJECE-V13I7P109

Enhancing Continuous Sign Language Recognition through MGPT-based Segmentation and Structured Position-Aware Decoding


Chauhan Pareshbhai Mansangbhai, Dineshkumar B. Vaghela, Mahesh M. Goyani, Udesang K. Jaliya

Received Revised Accepted Published
16 May 2026 11 Jun 2026 30 Jun 2026 29 Jul 2026

Citation :

Chauhan Pareshbhai Mansangbhai, Dineshkumar B. Vaghela, Mahesh M. Goyani, Udesang K. Jaliya, "Enhancing Continuous Sign Language Recognition through MGPT-based Segmentation and Structured Position-Aware Decoding," International Journal of Electronics and Communication Engineering, vol. 13, no. 7, pp. 127-139, 2026. Crossref, https://doi.org/10.14445/23488549/IJECE-V13I7P109

Abstract

Continuous Sign Language Recognition (CSLR) has always stood quite difficult due to issues like coarticulation effects, availability of weak temporal annotations, and large variability among different signers, especially in a limited-resource scenario such as Indian Sign Language (ISL). Most of the current methods use end-to-end sequence modeling, which leaves out the explicit temporal structure and does not work well with weak supervision. Here, we propose a structured CSLR system that combines motion-guided segmentation, multimodal representation learning, and position-aware decoding aimed at overcoming these issues. More importantly, we present an MGPT-based temporal segmentation method that uses optical-flow-driven motion signals and Gaussian peak modeling to separate continuous signing sequences into consistent motion segments, which results in the reduction of transitional ambiguity. The spatial-temporal features are obtained with the help of a dual-stream architecture that integrates ResNet50-based visual representations and skeleton keypoint features, being then temporally modeled by a multi-layer LSTM network. To improve sequence-level consistency, we also introduce a Word Position Graph (WPG) for structured decoding along with Gaussian-weighted frame voting to highlight informative temporal regions and, at the same time, downplay noisy transitions. The approach we suggested was tested on the ISL-CSLRT dataset with weak sentence-level supervision. The experimental results show that our framework reaches 92% accuracy and a Word Error Rate (WER) of 0.07, greatly beating the baseline voting strategies. Statistical verifications, including multi-run evaluation and significance testing, have confirmed the robustness of the improvements. Also, comparing with representative CSLR methods has shown that the method of explicit temporal segmentation and position-aware decoding is very effective, especially when the dataset is scarce. Besides, the results indicate that introducing motion-consistent segmentation and structured decision fusion seems to be a good way for updating the CSLR systems beyond simply endwise paradigms.

Keywords

Continuous Sign Language Recognition (CSLR), Indian Sign Language (ISL-CSLRT), MGPT-Based Segmentation, Motion-Guided Video Segmentation, Skeleton keypoint representation, ResNet50–LSTM, Word Position Graph (WPG), Gaussian-Weighted frame voting, Weak Sentence-Level annotation.

References

  1. Necati Cihan Camgöz et al., “Sign Language Transformers: Joint End-to-End Sign Language Recognition and Translation,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, pp. 10023-10033, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  2. Oscar Koller, “Quantitative Survey of the State of the Art in Sign Language Recognition,” arXiv preprint, pp. 1-40, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  3. Hao Zhou et al., “Spatial-Temporal Multi-Cue Network for Continuous Sign Language Recognition,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 7, pp. 13009-13016, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  4. Yuecong Min et al., “Visual Alignment Constraint for Continuous Sign Language Recognition,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, pp. 11522-11531, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  5. Dongxu Li et al., “Word-Level Deep Sign Language Recognition from Video: A New Large-Scale Dataset and Methods Comparison,” 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), Snowmass, CO, USA, pp. 1879-1889, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  6. Yao Du et al., “Full Transformer Network with Masking Future for Word-Level Sign Language Recognition,” Neurocomputing, vol. 500, pp. 115-123, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  7. Runpeng Cui, Hu Liu, and Changshui Zhang, “A Deep Neural Framework for Continuous Sign Language Recognition by Iterative Training,” IEEE Transactions on Multimedia, vol. 21, no. 7, pp. 1880-1891, 2019.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  8. Razieh Rastgoo, Kourosh Kiani, and Sergio Escalera, “Sign Language Recognition: A Deep Survey,” Expert Systems with Applications, vol. 164, pp. 1-89, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  9. Songyao Jiang et al., “Skeleton Aware Multi-modal Sign Language Recognition,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Nashville, TN, USA, pp. 328-338, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  10. Ashish Vaswani et al., “Attention is all you Need,” Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 1-11, 2017.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  11. Junfu Pu, Wengang Zhou, and Houqiang Li, “Iterative Alignment Network for Continuous Sign Language Recognition,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, pp. 4160-4169, 2019.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  12. Hannah Bull et al., “Aligning Subtitles in Sign Language Videos,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, pp. 11532-11541, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  13. Sijie Yang, Yuanjun Xiong, and Dahua Lin, “Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, pp. 7444-7452, 2018.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  14. Zhe Cao et al., “OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 1, pp. 172-186, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  15. Fan Zhang et al., “MediaPipe Hands: On-device Real-time Hand Tracking,” arXiv preprint, pp. 1-5, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  16. Pan Xie et al., “Multi-Scale Local-Temporal Similarity Fusion for Continuous Sign Language Recognition,” Pattern Recognition, vol. 136, pp. 1-33, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  17. Borui Miao et al., “DC-BVM: Dual-Channel Information Fusion Network based on Voting Mechanism,” Biomedical Signal Processing and Control, vol. 94, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  18. Hezhen Hu et al., “SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, pp. 11067-11076, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  19. Lu Meng, and Ronghui Li, “An Attention-Enhanced Multi-Scale and Dual Sign Language Recognition Network Based on a Graph Convolution Network,” Sensors, vol. 21, no. 4, pp. 1-22, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  20. Weichao Zhao et al., “Self-Supervised Representation Learning With Spatial-Temporal Consistency for Sign Language Recognition,” IEEE Transactions on Image Processing, vol. 33, pp. 4188-4201, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  21. R. Elakkiya, and B. Natarajan, “ISL-CSLTR: Indian Sign Language Dataset for Continuous Sign Language Translation and Recognition,” Mendeley Data V1, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  22. M. Madhiarasan, and Partha Pratim Roy, “A Comprehensive Review of Sign Language Recognition: Different Types, Modalities, and Datasets,” arXiv preprint, pp. 1-30, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  23. Mohammed Waleed Kadous, “Machine Recognition of Auslan Signs Using PowerGloves: Towards Large-Lexicon Recognition of Sign Language,” Proceedings of the Workshop on the Integration of Gesture in Language and Speech, pp. 1-10, 1996.
    [
    Google Scholar]
  24. C. Vogler, and D. Metaxas, “Parallel Hidden Markov Models for American Sign Language Recognition,” Proceedings of the Seventh IEEE International Conference on Computer Vision, Kerkyra, Greece, pp. 116-122, 1999.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  25. Helen Cooper, Brian Holt, and Richard Bowden, Sign Language Recognition, Visual Analysis of Humans, pp. 539-562, 2011.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  26. Sarah Alyami, and Hamzah Luqma, “Swin-MSTP: Swin Transformer with Multi-Scale Temporal Perception for Continuous Sign Language Recognition,” Neurocomputing, vol. 617, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]