20 papers
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
Fan Zhang, Guangming Yao, Jinyang Wu +6
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approx…
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
Jiahao Huang, Zheng Lian, Jingyi Zhang +3
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific sp…
AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning
Xiaojiang Peng, Kai Peng, Jie Lu +3
Vision-Language-Action (VLA) models excel at end-to-end robotic manipulation but struggle with out-of-distribution (OOD) generalization when familiar sub-tasks are recombined in un…
OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing
Zebang Cheng, Shuimu Chen, Boxue Yang +7
Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounce…
RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI
Kuofei Fang, Xinyi Che, Haomin Ouyang +12
Embodied AI is a prominent research topic in both academia and industry. Current research centers on completing tasks based on explicit user instructions. However, for robots to in…
Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking
Fan Zhang, Vireo Zhang, Shengju Qian +7
Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with recent efforts shifting from…