10 papers · 1 filter
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Haitao Lin, Hanyang Yu, Jingshun Huang +5
Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-spec…
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
Changli Wu, Haodong Wang, Jiayi Ji +5
Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with o…
Mind-of-Director: Multi-modal Agent-Driven Film Previsualization via Collaborative Decision-Making
Shufeng Nan, Mengtian Li, Sixiao Zheng +3
We present Mind-of-Director, a multi-modal agent-driven framework for film previz that models the collaborative decision-making process of a film production team. Given a creative…
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
Sixiao Zheng, Minghao Yin, Wenbo Hu +3
Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as vi…
ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
Yifan Li, Yingda Yin, Lingting Zhu +4
Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances.…
Enhancing Video Inpainting with Aligned Frame Interval Guidance
Ming Xie, Junqiu Yu, Qiaole Dong +2
Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. N…