7 papers · 1 filter
Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs
Wenzhang Sun, Chunfeng Wang, Xiangchen Yin +3
Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level.…
Holo-World: Unified Camera, Object and Weather Control for Video World Model
Xiangchen Yin, Wenzhang Sun, Jiahui Yuan +6
Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls r…
Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
Zhangchi Hu, Wenzhang Sun, Xiangchen Yin +5
Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions while synthesizing disoccluded…
MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration
Wenzhang Sun, Zhenyu Wang, Zhangchi Hu +3
Generating long-form audio-visual stories from a short user prompt remains challenging due to an intent-execution gap, where high-level narrative intent must be preserved across co…
DrivingScene: A Multi-Task Online Feed-Forward 3D Gaussian Splatting Method for Dynamic Driving Scenes
Qirui Hou, Wenzhang Sun, Chang Zeng +3
Real-time, high-fidelity reconstruction of dynamic driving scenes is challenged by complex dynamics and sparse views, with prior methods struggling to balance quality and efficienc…
PAGS: Priority-Adaptive Gaussian Splatting for Dynamic Driving Scenes
Ying A, Wenzhang Sun, Chang Zeng +3
Reconstructing dynamic 3D urban scenes is crucial for autonomous driving, yet current methods face a stark trade-off between fidelity and computational cost. This inefficiency stem…