• Home
  • Search
  • Context-Aware Multimodal Emotion Recognition
  • Cite Icon6
  • https://doi.org/10.1007/978-981-16-7618-5_5Copy DOI Icon

Context-Aware Multimodal Emotion Recognition

  • Jan 1, 2022
  • Aaishwarya Khalane +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Abstract Making human–computer interaction more organic and personalized for users essentially demands advancement in human emotion recognition. Emotions are perceived by humans considering multiple factors such as facial expressions, voice tonality, and information context. Although significant research has been conducted in the area of unimodal/multimodal emotion recognition in videos using acoustic/visual features, few papers have explored the potential of textual information obtained from the video utterances. Humans experience emotions through their audio-visual and linguistic senses, making it quintessential to take the latter into account. This paper outlines two different algorithms for recognizing multimodal emotional expressions in online videos. In addition to acoustic (speech), visual (facial), and textual (utterances) feature extraction using BERT, we utilize bidirectional LSTMs to capture the context between utterances. To obtain richer sequential information, we also implement a multi-head self-attention mechanism. Our analysis utilizes the benchmarking CMU multimodal opinion sentiment and emotion intensity (CMU-MOSEI) dataset, which is the largest dataset for sentiment analysis and emotion recognition to date. Our experiments result in improved F1 scores in comparison to the baseline models.KeywordMultimodalEmotionRecognitionCMU-MOSEIMulti-head attentionBERTContext-aware

Similar Papers
  • Conference Article
  • Citations26

Emotion recognition from facial expressions for 3D videos using siamese network

  • Jun 16, 2021
  • Divina Lawrance +1
  • Conference Article

Development of Video-Based Emotion Recognition System using Transfer Learning

  • Sep 26, 2022
  • Teddy Surya Gunawan +5
  • Research Article

Enhancing emotion recognition in virtual reality: a multimodal dataset and a temporal emotion detector

  • Nov 24, 2025
  • Frontiers in Psychology
  • Chenxin Qu +7
  • Dissertation

Towards context information estimation for wearables using minimal sensors

  • Jan 01, 2025
  • Mohammad Rahmani
  • Research Article
  • Citations34

Recognition of facial and musical emotions in Parkinson's disease

  • Dec 24, 2012
  • European Journal of Neurology
  • A Saenz +5
  • PDF
  • Research Article
  • Citations16

Feature Extraction Network with Attention Mechanism for Data Enhancement and Recombination Fusion for Multimodal Sentiment Analysis

  • Aug 24, 2021
  • Information
  • Qingfu Qi +2
  • Research Article
  • Citations174

A survey of emotion recognition methods with emphasis on E-Learning environments

  • Aug 27, 2019
  • Journal of Network and Computer Applications
  • Maryam Imani +1
  • Research Article

Improved Attention-Enhanced Efficient Face-Transformer Model for Multimodal Elderly Emotion Recognition in Smart Homes

  • Aug 26, 2025
  • Informatica
  • Peiyao Li
  • PDF
  • Research Article
  • Citations11

Multimodal interaction enhanced representation learning for video emotion recognition.

  • Dec 19, 2022
  • Frontiers in Neuroscience
  • Xiaohan Xia +2
  • Research Article
  • Citations3

Multimodal information fusion method in emotion recognition in the background of artificial intelligence

  • Mar 12, 2024
  • Internet Technology Letters
  • Zhen Dai +2
  • Research Article
  • Citations2

Understanding basic and social emotions in Alzheimer's disease and frontotemporal dementia

  • Feb 07, 2025
  • Frontiers in Psychology
  • Carlotta Sola +9
  • Research Article

Novel Deep Learning Methods for Identifying Human Emotions: A Comprehensive Review

  • Jul 01, 2025
  • Journal of Science Engineering Technology and Management Sciences
  • Ms Barkha Sain +2
  • Research Article
  • Citations267

Multimodal emotion recognition in speech-based interaction using facial expression, body gesture and acoustic analysis

  • Dec 12, 2009
  • Journal on Multimodal User Interfaces
  • Loic Kessous +2
  • Conference Article
  • Citations38

Automatic Group Level Affect and Cohesion Prediction in Videos

  • Sep 01, 2019
  • Garima Sharma +2
  • Research Article
  • Citations16

Automatic highlights extraction for drama video using music emotion and human face features

  • Jan 03, 2013
  • Neurocomputing
  • Keng-Sheng Lin +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.