14 papers
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
Daiqing Wu, Dongbao Yang, Jiashu Yao +4
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intellige…
Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms
Jiashu Yao, Heyan Huang, Daiqing Wu +5
GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications violate this assumption, as their…
Multimodal Emotion Recognition with Large Language Models
Hongrui Zhang, Daiqing Wu, Yangyang Li +4
Multimodal Emotion Recognition (MER) focuses on identifying and interpreting emotions from modality-compound inputs. Closely mirroring human cognitive processes in real-world envir…
Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
Chenmin Yu, Liu Yu, Daiqing Wu +3
Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored,…
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
Yan Zhang, Daiqing Wu, Huawen Shen +2
Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capability for autonomous GUI agents.…
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
Daiqing Wu, Xuan Zhang, Dongbao Yang +7
The maturation of Large Audio Language Models (LALMs) has raised growing expectations for them to comprehend complex audio much like humans. Current efforts primarily replicate tex…