2 papers
cs.CV2025
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition
Xudong Yang, Yizhang Zhu, Hanfeng Liu +3
Conventional Multi-modal multi-label emotion recognition (MMER) assumes complete access to visual, textual, and acoustic modalities. However, real-world multi-party settings often…
cs.CV2024
AskChart: Universal Chart Understanding through Textual Enhancement
Xudong Yang, Yifan Wu, Yizhang Zhu +2
Chart understanding tasks such as ChartQA and Chart-to-Text involve automatically extracting and interpreting key information from charts, enabling users to query or convert visual…