Preview of the new IC2 website. It is not public yet and is hidden from search engines.

Publications

Cross-Modal Emotion Discrepancy Detection: Detecting When Voice Reveals Emotions That Text Conceals

KH Hein, JH Chan

Emotion Recognition and Brain InformaticsWeb Intelligence

Abstract

Traditional multi-modal emotion recognition systems fuse text and audio to improve a single emotion prediction by seeking inter-modal agreement. This work takes the opposite approach: we model the disagreement between text sentiment and voice prosody as a signal for emotional discrepancy—an approach we term asymmetric fusion. Using the RAVDESS dataset (2,880 recordings from 24 actors), where semantically neutral text is paired with emotionally varied speech, we extract 26 features across three categories: text features from DistilRoBERTa, voice prosodic features from librosa and wav2vec2-based speech emotion recognition, and cross-modal discrepancy features. We systematically compare eight classifiers—including logistic regression, SVM, Random Forest, XGBoost, LightGBM, and Gradient Boosting—using stratified grouped 5-fold cross-validation. Gradient Boosting achieves the best performance (F1 = 0.954, AUC = 0.966), while logistic regression provides an interpretable alternative (F1 = 0.918, AUC = 0.947). A feature selection study reveals that the top-10 features match or exceed full-set performance, and per-emotion analysis shows perfect detection for angry, disgust, and surprised categories. Cross-dataset validation on MELD (13,079 utterances with naturally varied text) confirms that cross-modal features contribute substantially when text content varies (mean F1 improvement of +0.160 over voice-only), compared to +0.001 on RAVDESS where text is uniform, validating the asymmetric fusion approach. Feature importance analysis identifies voice energy variation as the dominant predictor, demonstrating the viability of interpretable, feature-engineered models for multi-modal emotion analysis. Code and resources are publicly available at https://github.com/PhilixTheExplorer/cross-modal-emotion-discrepancy.

Authors: Kaung Hset Hein, Jonathan H. Chan

DOI · Google Scholar