10 papers
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Xu Guo, Zhengxuan Wei, Xinghui Li +11
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…
Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding
Shiyu Li, Zhiyuan Hu, Yifan Wang +3
Omni-modal retrieval promises a single embedding space for text, image, video, document, and audio inputs, but building such a unified retriever is difficult since these modalities…
Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding
Yifan Wang, Peiming Li, Shiyu Li +5
While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent traini…
VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling
Xunzhi Xiang, Zixuan Duan, Yabo Chen +8
Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually e…
SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models
Zhengxuan Wei, Yi Dong, Zonghui Li +6
Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, existing LoRA merging techniq…
Geometry-Aware Implicit Memory for Video World Models
Zhengxuan Wei, Xu Guo, Xinghui Li +8
Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window…