Abstract
Multimodal large language models (MLLMs) have demonstrated strong performance across diverse multimodal tasks, achieving promising outcomes. However, their application to emotion recognition in natural images remains underexplored. MLLMs struggle to handle ambiguous emotional expressions and implicit affective cues, whose capability is crucial for affective understanding but largely overlooked. To address these challenges, we propose MERMAID, a novel multi-agent framework that integrates a multi-perspective self-reflection module, an emotion-guided visual augmentation module, and a cross-modal verification module. These components enable agents to interact across modalities and reinforce subtle emotional semantics, thereby enhancing emotion recognition and supporting autonomous performance. Extensive experiments show that MERMAID outperforms existing methods, achieving absolute accuracy gains of 8.70%-27.90% across diverse benchmarks and exhibiting greater robustness in emotionally diverse scenarios.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing |
| Editors | Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng |
| Publisher | Association for Computational Linguistics |
| Pages | 24639-24655 |
| Number of pages | 17 |
| ISBN (Electronic) | 9798891763326 |
| DOIs | |
| Publication status | Published - Nov 2025 |
| Event | 30th Conference on Empirical Methods in Natural Language Processing 2025 - Suzhou, China Duration: 4 Nov 2025 → 9 Nov 2025 |
Conference
| Conference | 30th Conference on Empirical Methods in Natural Language Processing 2025 |
|---|---|
| Abbreviated title | EMNLP 2025 |
| Country/Territory | China |
| City | Suzhou |
| Period | 4/11/25 → 9/11/25 |
ASJC Scopus subject areas
- Computational Theory and Mathematics
- Computer Science Applications
- Information Systems
- Linguistics and Language
Fingerprint
Dive into the research topics of 'MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion Recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver