4 papers
EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks
Tongtong Feng, Xin Wang, Zekai Zhou +5
Completing Long-Horizon (LH) tasks in open-ended worlds is an important yet difficult problem for embodied agents. Existing approaches suffer from two key challenges: (1) they heav…
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
Guangyao Li, Xin Wang, Wenwu Zhu
When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing,…
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
Houlun Chen, Xin Wang, Guangyao Li +4
Long video understanding is challenging due to rich and complicated multimodal clues in long temporal range.Current methods adopt reasoning to improve the model's ability to analyz…
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
Yu-Wei Zhan, Xin Wang, Hong Chen +6
Video Large Language Models (Video LLMs) have shown impressive performance across a wide range of video-language tasks. However, they often fail in scenarios requiring a deeper und…