9 papers
PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians
Jiang Qin, Chunji Lv, Yangguang Wei +6
Efficient, fully automatic, and physically plausible 4D Gaussian synthesis is an important goal for dynamic scene generation. Recent physics-based methods couple 3D Gaussians with…
Physically Plausible Video Generation via Visual-Semantic Chain-of-Events Conditioning
Zixuan Wang, Yixin Hu, Wen Li +4
Physically Plausible Video Generation (PPVG) seeks to synthesize videos consistent with physical principles, yet remains challenging due to underspecified natural language conditio…
Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
Bo Zhang, Wenxin Wang, Feng Chen +4
Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting query-relevant frames. How…
Agentic Collaborative Cognition for Zero-Shot 3D Understanding
Wenxin Wang, Bo Zhang, Feng Chen +4
Recent advancements have explored agentic zero-shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language Models (MLLMs). However, exi…
PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models
Xianghui Wang, Feng Chen, Wenbo Zhang +4
Vision-Language-Action (VLA) models provide a unified paradigm for robotic manipulation, yet their real-world deployment is often bottlenecked by execution efficiency. While existi…
Training-free Motion Factorization for Compositional Video Generation
Zixuan Wang, Ziqin Zhou, Feng Chen +4
Compositional video generation aims to synthesize multiple instances with diverse appearance and motion. However, current approaches mainly focus on binding semantics, neglecting t…