Call For Paper - Upcoming Conferences

Research Article | Open Access | Download PDF
Volume 13 | Issue 7 | Year 2026 | Article Id. IJECE-V13I7P102 | DOI : https://doi.org/10.14445/23488549/IJECE-V13I7P102

An Attention-Guided Framework for Feature-Level and Decision-Level Fusion in Multimodal Emotion Recognition


Chintan Chatterjee, Brijesh Bhatt

Received Revised Accepted Published
11 Apr 2026 30 May 2026 17 Jun 2026 29 Jul 2026

Citation :

Chintan Chatterjee, Brijesh Bhatt, "An Attention-Guided Framework for Feature-Level and Decision-Level Fusion in Multimodal Emotion Recognition," International Journal of Electronics and Communication Engineering, vol. 13, no. 7, pp. 17-34, 2026. Crossref, https://doi.org/10.14445/23488549/IJECE-V13I7P102

Abstract

The Multimodal Emotion Recognition (MER) is a critical aspect in the development of human-computer interaction since it is a synthesis of non-redundant information based on the use of text, audio, and visual modalities. However, the comparative effectiveness of disparate fusion strategies and mechanisms of attention in MER has not been thoroughly studied on a single experimental paradigm. In line with this, the present study engages in a comparative analysis of four multimodal configurations, namely: early Fusion without attention, early Fusion with attention, late Fusion without attention, and late Fusion complemented by attention. Each of these configurations is evaluated on the Multimodal Emotion Lines Dataset (MELD) using harmonized training protocols to ensure a fair comparison. The methodologies use pre-trained architectures of BERT to support textual representation, Wav2Vec to support acoustic encoding, and TimesFormer to support visual streams to extract features that are modality-specific. These characteristics are then fused through the above fusion tactics. Attention modules that constitute cross attention, hierarchical attention, and self-attention modules are incorporated to enable cross-modal as well as intra-modal features interaction. The results of empirical studies have shown that there are recognizable differences between the models, where attention-enhanced schemes have better measures of performance, especially improved in terms of accuracy and weighted F1-score. However, they still have residual problems, including the imbalance of the classes that influence minority emotion categories. Such findings explain the impact created by specific fusion paradigms and attention structure on the multimodal emotion recognition performance. In turn, this research provides a methodological framework that can be used to develop more effective and understandable MER systems using a systematic and structured comparative framework.

Keywords

Attention mechanisms, Early Fusion, Late Fusion, Multimodal Emotion Recognition, Multimodal Fusion.

References

  1. Qingping Zhou, “HGLER: A Hierarchical Heterogeneous Graph Networks for Enhanced Multimodal Emotion Recognition in Conversations,” PLoS One, vol. 20, no. 9, pp. 1-19, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  2. Yuntao Shou et al., “A Comprehensive Survey on Multi-Modal Conversational Emotion Recognition with Deep Learning,” ACM Transactions on Information Systems, vol. 44, no. 2, pp. 1-48, 2026.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  3. Wei Dai et al., “A Novel Approach for Multimodal Emotion Recognition: Multimodal Semantic Information Fusion,” arXiv preprint, pp. 1-13, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  4. Yuanyuan Sun, and Ting Zhou, “DialogueMLLM: Transforming Multimodal Emotion Recognition in Conversation through Instruction-Tuned MLLM,” IEEE Access, vol. 13, pp. 121048-121060, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  5. Chhavi Dixit, and Shashank Mouli Satapathy, “Deep CNN with Late Fusion for Real Time Multimodal Emotion Recognition,” Expert Systems with Applications, vol. 240, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  6. Yong Zhang, Cheng Cheng, and YiDie Zhang, “Multimodal Emotion Recognition based on Manifold Learning and Convolution Neural Network,” Multimedia Tools and Applications, vol. 81, pp. 33253-33268, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  7. Jiarong He, “A Multimodal Approach for Emotion Recognition in Conversations Using the MELD Dataset,” 2025 Asia-Europe Conference on Cybersecurity, Internet of Things and Soft Computing (CITSC), Rimini, Italy, pp. 54-58, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  8. Yun Liu et al., “Toward Multimodal Sentiment Analysis with a Self-Supervised Knowledge-Augmented Network,” Expert Systems with Applications, vol. 314, 2026.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  9. Wuzhen Shi et al., “Identity and Modality Attributes Driven Multimodal Fusion Networks for Emotion Recognition in Conversations,” IEEE Transactions on Multimedia, vol. 27, pp. 4361-4371, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  10. Shiyun Zhao, Jinchang Ren, and Xiaojuan Zhou, “Cross-modal Gated Feature Enhancement for Multimodal Emotion Recognition in Conversations,” Scientific Reports, vol. 15, pp. 1-13, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  11. Gyanendra K. Verma, and Uma Shanker Tiwary, “Multimodal Fusion Framework: A Multiresolution Approach for Emotion Classification and Recognition from Physiological Signals,” NeuroImage, vol. 102, no. 1, pp. 162-172, 2014.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  12. Komal Anadkat et al., “Enhancing Emotion Recognition with Multimodel Approach using Deep Neural Networks,” Reliability: Theory & Applications, vol. 20, no. 1, pp. 1-13, 2025.
    [
    Google Scholar] [Publisher Link]
  13. Samuel Kakuba, Alwin Poulose, and Dong Seog Han, “Deep Learning Approaches for Bimodal Speech Emotion Recognition: Advancements, Challenges, and a Multi-Learning Model,” IEEE Access, vol. 11, pp. 113769-113789, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  14. Mingyu, Zhou Jiawei, and Wei Ning, “AFR-BERT: Attention-based Mechanism Feature Relevance Fusion Multimodal Sentiment Analysis Model,” PLoS One, vol. 17, no. 9, pp. 1-20, 2022.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  15. Licai Sun et al., “MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression Recognition,” Proceedings of the 31st ACM International Conference on Multimedia, pp. 6110-6121, 2023.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  16. Alexei Baevski et al., “Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” arXiv preprint, pp. 1-19, 2020.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  17. Gedas Bertasius, Heng Wang, and Lorenzo Torresani, “Is Space-Time Attention All you Need for Video Understanding?,” arXiv Preprint, pp. 1-13, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  18. Cheng Cheng et al., “Dense Graph Convolutional with Joint Cross-Attention Network for Multimodal Emotion Recognition,” IEEE Transactions on Computational Social Systems, vol. 11, no. 5, pp. 6672-6683, 2024.
    [CrossRef] [Google Scholar] [Publisher Link]
  19. Yunhong Liao et al., “Attention-Driven Adaptive Deep Canonical Correlation Analysis for Multimodal Sentiment Analysis with EEG and Eye Movement Data,” 2025 IEEE World AI IoT Congress (AIIoT), Seattle, WA, USA, pp. 409-415, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  20. Mohammad Soleymani, Maja Pantic, and Thierry Pun, “Multimodal Emotion Recognition in Response to Videos,” IEEE Transactions on Affective Computing, vol. 3, no. 2, pp. 211-223, 2012.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  21. Wei-Bang Jiang et al., “SEED-VII: A Multimodal Dataset of Six Basic Emotions with Continuous Labels for Emotion Recognition,” IEEE Transactions on Affective Computing, vol. 16, no. 2, pp. 969-985, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  22. Nicu Sebe, Ira Cohen, and Thomas S. Huang, Multimodal Emotion Recognition, Handbook of Pattern Recognition and Computer Vision, 4th ed., pp. 387-409, 2005.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  23. Panagiotis Tzirakis et al., “End-to-End Multimodal Emotion Recognition Using Deep Neural Networks,” IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 8, pp. 1301-1309, 2017.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  24. Wei Liu, Wei-Long Zheng, and Bao-Liang Lu, “Emotion Recognition using Multimodal Deep Learning,” 23rd International Conference Neural Information Processing, Kyoto, Japan, vol. 9948, pp. 521-529, 2016.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  25. You Wu, Qingwei Mi, and Tianhan Gao, “A Comprehensive Review of Multimodal Emotion Recognition: Techniques, Challenges, and Future Directions,” Biomimetics, vol. 10, no. 7, pp. 1-24, 2025.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  26. Huan Liu et al., “EEG-Based Multimodal Emotion Recognition: A Machine Learning Perspective,” IEEE Transactions on Instrumentation and Measurement, vol. 73, pp. 1-29, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  27. Wei-Long Zheng et al., “EmotionMeter: A Multimodal Framework for Recognizing Human Emotions,” IEEE Transactions on Cybernetics, vol. 49, no. 3, pp. 1110-1122, 2019.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  28. Zebang Cheng et al., “Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning,” arXiv preprint, pp. 110805-110853, 2024.
    [CrossRef] [Google Scholar] [Publisher Link]
  29. Samira Hazmoune, and Fateh Bougamouza, “Using Transformers for Multimodal Emotion Recognition: Taxonomies and State of the Art Review,” Engineering Applications of Artificial Intelligence, vol. 133, 2024.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  30. Cristina Luna-Jiménez et al., “Multimodal Emotion Recognition on Ravdess Dataset using Transfer Learning,” Sensors, vol. 21, no. 22, pp. 1-29, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  31. Deepanway Ghosal et al., “DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, pp. 154-164, 2019.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  32. Dou Hu et al., “DialogueCRN: Contextual Reasoning Networks for Emotion Recognition in Conversations,” Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, vol. 1, pp. 7042-7052, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]
  33. Jingwen Hu et al., “MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in Conversation,” Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, vol. 1, pp. 5666-5675, 2021.
    [
    CrossRef] [Google Scholar] [Publisher Link]