dorsal/arxiv
View SchemaA Unified Framework for Emotion Recognition and Sentiment Analysis via Expert-Guided Multimodal Fusion with Large Language Models
| Authors | Jiaqi Qiao, Xiujuan Xu, Xinran Li, Yu Liu |
|---|---|
| Categories | |
| ArXiv ID | 2601.07565vv1 |
| URL | https://arxiv.org/abs/2601.07565 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided multimodal fusion with large language models. Our approach features three specialized expert networks--a fine-grained local expert for subtle emotional nuances, a semantic correlation expert for cross-modal relationships, and a global context expert for long-range dependencies--adaptively integrated through hierarchical dynamic gating for context-aware feature selection. Enhanced multimodal representations are integrated with LLMs via pseudo token injection and prompt-based conditioning, enabling a single generative framework to handle both classification and regression through natural language generation. We employ LoRA fine-tuning for computational efficiency. Experiments on bilingual benchmarks (MELD, CHERMA, MOSEI, SIMS-V2) demonstrate consistent improvements over state-of-the-art methods, with superior cross-lingual robustness revealing universal patterns in multimodal emotional expressions across English and Chinese. We will release the source code publicly.
{
"annotation_id": "053aca06-8a9e-40a2-af50-c0df2a93515c",
"date_created": "2026-02-17T05:53:11.886000Z",
"date_modified": "2026-02-17T05:53:11.886000Z",
"file_hash": "0dbce122f55646e156fca9b1213da789042af2c1da2a3be3f3a2f10d8634e3a0",
"private": false,
"record": {
"abstract": "Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided multimodal fusion with large language models. Our approach features three specialized expert networks--a fine-grained local expert for subtle emotional nuances, a semantic correlation expert for cross-modal relationships, and a global context expert for long-range dependencies--adaptively integrated through hierarchical dynamic gating for context-aware feature selection. Enhanced multimodal representations are integrated with LLMs via pseudo token injection and prompt-based conditioning, enabling a single generative framework to handle both classification and regression through natural language generation. We employ LoRA fine-tuning for computational efficiency. Experiments on bilingual benchmarks (MELD, CHERMA, MOSEI, SIMS-V2) demonstrate consistent improvements over state-of-the-art methods, with superior cross-lingual robustness revealing universal patterns in multimodal emotional expressions across English and Chinese. We will release the source code publicly.",
"arxiv_id": "2601.07565",
"authors": [
"Jiaqi Qiao",
"Xiujuan Xu",
"Xinran Li",
"Yu Liu"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "A Unified Framework for Emotion Recognition and Sentiment Analysis via Expert-Guided Multimodal Fusion with Large Language Models",
"url": "https://arxiv.org/abs/2601.07565",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "13edebdf-dffc-425d-829c-f69c922c22db",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}