Call For Paper September 2026

Research Article | Open Access | Download PDF
Volume 13 | Issue 9 | Year 2026 | Article Id. IJECE-V13I9P113 | DOI : https://doi.org/10.14445/23488549/IJECE-V13I9P113

ECAPA-TF: A Hybrid ECAPA-Transformer Framework for Multilingual Speaker Identification using the NISP Corpus


Muliya Pritibahen Jagjivanbhai, Ashwin Patni

Received Revised Accepted Published
30 May 2026 07 Aug 2026 19 Aug 2026 29 Sep 2026

Citation :

Muliya Pritibahen Jagjivanbhai, Ashwin Patni, "ECAPA-TF: A Hybrid ECAPA-Transformer Framework for Multilingual Speaker Identification using the NISP Corpus," International Journal of Electronics and Communication Engineering, vol. 13, no. 9, pp. 208-223, 2026. Crossref, https://doi.org/10.14445/23488549/IJECE-V13I9P113

Abstract

Speaker identification in multiple languages has received increasing attention in speech processing because of the growing use of voice-based security applications in multilingual settings. Traditional speaker identification systems have encountered difficulties ensuring reliable performance in bilingual and cross-language speech scenarios due to their concentration on local acoustic models and neglecting long-range contextual information. This research presents a novel hybrid architecture referred to as ECAPA-TF, which combines ECAPA-TDNN-based local acoustic modelling with transformer-based global contextual modelling for multilingual speaker identification. The proposed model is trained using the Native and Indian Speakers of Indian Phonetics (NISP) database consisting of 345 bilingual speakers of India, who are proficient in both English and native Indian languages, including Malayalam, Hindi, Kannada, Telugu and Tamil. The first step involves extracting Mel-Frequency Cepstral Coefficient (MFCC) feature vectors (40-dimensional) from the speech signal and normalizing the data prior to applying ECAPA residual blocks equipped with squeeze-excitation attention. Subsequently, the Transformer self-attention mechanism is implemented to identify long-range temporal and contextual dependencies in speech data. Attention pooling and Additive Angular Margin SoftMax (AAM-SoftMax) are also leveraged to create discriminative speaker embeddings. Performance is evaluated using Accuracy, recall, precision, confusion matrix, F1-score and ROC-AUC. The developed ECAPA-TF model achieved an accuracy of 92.93%, precision of 93.64%, recall of 93.02%, macro F1-score of 92.94%, and an AUC of 0.965. Under the same NISP experimental protocol, the proposed model outperformed the internally implemented Deep 1D-CNN, CNN-LSTM, and ECAPA baseline models. Compared with the ECAPA baseline, ECAPA-TF improved accuracy by 1.24 percentage points and macro F1-score by 1.36 percentage points. The confusion matrix and per-speaker evaluation further validate the effectiveness and reliability of the proposed approach under various multilingual speech scenarios.

Keywords

Multilingual speaker identification, Deep learning, ECAPA-TF, MFCC, AAM-SoftMax, NISP Corpus.

References

  1. Juraj Kacur and Peter Truchly, “Acoustic and Auxiliary Speech Features for Speaker Identification System,” 2015 57th International Symposium ELMAR (ELMAR), Zadar, Croatia, pp. 109-112, 2015.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  2. Nourah M. Almarshady, Adal A. Alashban, and Yousef A. Alotaibi, “Analysis and Investigation of Speaker Identification Problems Using Deep Learning Networks and the YOHO English Speech Dataset,” Applied Sciences, vol. 13, no. 17, pp. 1-19, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  3. John H.L. Hansen and Taufiq Hasan, “Speaker Recognition by Machines and Humans: A Tutorial Review,” IEEE Signal Processing Magazine, vol. 32, no. 6, pp. 74-99, 2015.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  4. Andreas Nautsch et al., “Preserving Privacy in Speaker and Speech Characterisation,” Computer Speech & Language, vol. 58, pp. 441-480, 2019.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  5. Ajan Ahmed and Masudul H. Imtiaz, “Quantifying the Relationship Between Speech Quality Metrics and Biometric Speaker Recognition Performance Under Acoustic Degradation,” Signals, vol. 7, no. 1, 2026.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  6. Ke-Ming Lyu, Ren-yuan Lyu, and Hsien-Tsung Chang, “Real-Time Multilingual Speech Recognition and Speaker Diarization System based on Whisper Segmentation,” PeerJ Computer Science, vol. 10, pp. 1-19, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  7. S. Pruzansky and M.V. Mathews, “Talker‐Recognition Procedure Based on Analysis of Variance,” The Journal of the Acoustical Society of America, vol. 35, no. 11, pp. 2041-2047, 1964.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  8. S. Furui, “Cepstral Analysis Technique for Automatic Speaker Verification,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 29, no. 2, pp. 254-272, 1981.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  9. S. Davis, and P. Mermelstein, “Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 28, no. 4, pp. 357-366, 1980.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  10. Hynek Hermansky, “Perceptual Linear Predictive (PLP) Analysis of Speech,” The Journal of the Acoustical Society of America, vol. 87, no. 4, pp. 1738-1752, 1990.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  11. Douglas A. Reynolds, Thomas F. Quatieri, and Robert B. Dunn, “Speaker Verification Using Adapted Gaussian Mixture Models,” Digital Signal Processing, vol. 10, no. 1, pp. 19-41, 2000.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  12. Kunal Thakur and Ramesh K. Bhukya, “Speaker Authentication Using GMM-UBM,” 2022 IEEE 9th Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON), Prayagraj, India, pp. 1-6, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  13. Mitchell McLaren, and David van Leeuwen, “Improved Speaker Recognition when Using I-Vectors from Multiple Speech Sources,” 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Prague, Czech Republic, pp. 5460-5463, 2011.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  14. Patrick Kenny et al., “Joint Factor Analysis versus Eigenchannels in Speaker Recognition,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 15, no. 4, pp. 1435-1447, 2007.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  15. Ehsan Variani et al., “Deep Neural Networks For Small Footprint Text-Dependent Speaker Verification,” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy, pp. 4052-4056, 2014.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  16. David Snyder et al., “X-Vectors: Robust DNN Embeddings for Speaker Recognition,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, pp. 5329-5333, 2018.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  17. David Snyder et al., “Deep Neural Network-Based Speaker Embeddings for End-to-End Speaker Verification,” 2016 IEEE Spoken Language Technology Workshop (SLT), San Diego, CA, USA, pp. 165-170, 2016.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  18. Rahul Vijaykumar et al., “Descriptor: Extended-Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD-SVDSR),” IEEE Data Descriptions, vol. 3, pp. 93-99, 2026.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  19. Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN-based Speaker Verification,” arXiv preprint, pp. 3830-3834, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  20. Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, “VoxCeleb : A Large-Scale Speaker Identification Dataset,” arXiv preprint, pp. 1-6, 2017.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  21. Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, “VoxCeleb2 : Deep Speaker Recognition,” arXiv preprint, pp. 1-6, 2018.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  22. Vassil Panayotov et al., “Librispeech: An ASR Corpus Based on Public Domain Audio Books,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia, pp. 5206-5210, 2015.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  23. Alexei Baevski et al., “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” Proceedings of the 34th International Conference on Neural Information Processing Systems, pp. 12449-12460, 2020.
    [
    Google Scholar] [Publisher Link]
  24. Wei-Ning Hsu et al., “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451-3460, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  25. Zhanibek Kozhirbayev et al., “Speaker Recognition for Robotic Control via an IoT Device,” 2018 World Automation Congress (WAC), Stevenson, WA, USA, pp. 1-5, 2018.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  26. Ayu Mawadda Warohma, Hilwadi Hindersah, and Dessi Puji Lestari, “Speaker Recognition Using MobileNetV3 for Voice-Based Robot Navigation,” 2024 11th International Conference on Advanced Informatics: Concept, Theory and Application (ICAICTA), Singapore, pp. 1-6, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  27. Jee-weon Jung et al., “RawNet: Advanced End-to-End Deep Neural Network using Raw Waveforms for Text-Independent Speaker Verification,” arXiv preprint, pp. 1-5, 2019.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  28. Jee-weon Jung et al., Pushing the Limits of Raw Waveform Speaker Recognition, arXiv preprint, pp. 1-5, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  29. Seo-Hyun Kim, Tae-Wan Kim, and Keun-Chang Kwak, “Speaker Recognition Based on the Combination of SincNet and Neuro-Fuzzy for Intelligent Home Service Robots,” Electronics, vol. 14, no. 18, pp. 1-22, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  30. Anuj Diwan et al., “Multilingual and Code-Switching ASR Challenges for Low-Resource Indian Languages,” arXiv preprint, pp. 1-6, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  31. Mahesh K. Singh et al., “Feature Extraction and Classification Technique Based Speaker Identification System of Indian Regional Accent,” Franklin Open, vol. 16, pp. 1-9, 2026.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  32. Kristiawan Nugroho et al., “Enhanced Indonesian Ethnic Speaker Recognition using Data Augmentation Deep Neural Network,” Journal of King Saud University Computer and Information Sciences, vol. 34, no. 7, pp. 4375-4384, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  33. Vajratiya Vajrobol et al., “Enhancing Speaker Identification in Low-Resource Multilingual Languages using Hybrid MFCC-Chroma STFT and Transformer Encoder,” Multimedia Tools and Applications, vol. 84, no. 35, pp. 44251-44285, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  34. Braveenan Sritharan and Uthayasanker Thayasivam, “Advancing Multilingual Speaker Identification and Verification for Indo-Aryan and Dravidian Languages,” Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages, pp. 67-73, 2025.
    [
    Google Scholar] [Publisher Link]
  35. Ruijie Tao et al., “Self-Supervised Speaker Recognition with Loss-Gated Learning,” ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, pp. 6142-6146, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  36. Wenzao Li et al., “TDNN Architecture with Efficient Channel Attention and Improved Residual Blocks for Accurate Speaker Recognition,” Scientific Reports, vol. 15, no. 1, pp. 1-13, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  37. Rafizah Mohd Hanifa, Khalid Isa, and Shamsul Mohamad, “Speaker Ethnic Identification for Continuous Speech in Malay Language Using Pitch and MFCC,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 19, no. 1, pp. 207-214, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  38. Muzamil Ahmed et al., “An Enhanced Deep Learning Approach for Speaker Diarization using TitaNet, MarbelNet and Time Delay Network,” Scientific Reports, vol. 15, pp. 1-15, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  39. Md. Iftekharul Alam Efat et al., “Identifying Optimised Speaker Identification Model using Hybrid GRU-CNN Feature Extraction Technique,” International Journal of Computational Vision and Robotics, vol. 12, no. 6, pp. 662-685, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  40. Weijun Pan et al., “The Speaker Identification Model for Air-Ground Communication Based on a Parallel Branch Architecture,” Applied Sciences, vol. 15, no. 6, pp. 1-21, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  41. Hossein Zeinali et al., “BUT System Description to VoxCeleb Speaker Recognition Challenge 2019,” arXiv preprint, pp. 4-7, 2019.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  42. Arsha Nagrani et al., “VOxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge,” arXiv preprint, pp. 1-8, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  43. A. Gusev et al., “SdSVC Challenge 2021 : Tips and Tricks to Boost the Short-duration Speaker Verification System Performance,” Proceedings of Interspeech 2021, pp. 2307-2311, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  44. Andrew Brown et al., “VoxSRC 2021 : The Third VoxCeleb Speaker Recognition Challenge,” arXiv preprint, pp. 1-8, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  45. Shareef Babu Kalluri et al., “NISP: A Multi-lingual Multi-accent Dataset for Speaker Profiling,” ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, pp. 6953-6957, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  46. Ashish Vaswani et al., “Attention is all you Need,” Advances in Neural Information Processing Systems, vol. 30, 2017.
    [
    Google Scholar] [Publisher Link]