15 papers
Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs
Wenzhang Sun, Chunfeng Wang, Xiangchen Yin +3
Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level.…
Holo-World: Unified Camera, Object and Weather Control for Video World Model
Xiangchen Yin, Wenzhang Sun, Jiahui Yuan +6
Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls r…
Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
Zhangchi Hu, Wenzhang Sun, Xiangchen Yin +5
Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions while synthesizing disoccluded…
RiO-DETR: DETR for Real-time Oriented Object Detection
Zhangchi Hu, Yifan Zhao, Yansong Peng +8
We present RiO-DETR: DETR for Real-time Oriented Object Detection, the first real-time oriented detection transformer to the best of our knowledge. Adapting DETR to oriented boundi…
MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration
Wenzhang Sun, Zhenyu Wang, Zhangchi Hu +3
Generating long-form audio-visual stories from a short user prompt remains challenging due to an intent-execution gap, where high-level narrative intent must be preserved across co…
DeCo-VAE: Learning Compact Latents for Video Reconstruction via Decoupled Representation
Xiangchen Yin, Jiahui Yuan, Zhangchi Hu +5
Existing video Variational Autoencoders (VAEs) generally overlook the similarity between frame contents, leading to redundant latent modeling. In this paper, we propose decoupled V…