22 papers · 1 filter
From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation
Zhefan Rao, Bin Zou, Xuanhua He +5
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-…
USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes
Li-Heng Chen, Haokai Pang, Chengye Su +7
Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-or…
MSEditor: Toward Consistent Multi-Shot Video Editing
Kunyu Feng, Yue Ma, Bingyuan Wang +6
In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos…
LiveLight: Real-time Streaming Video Relighting with Interactive Control
Yue Ma, Jiangming Wang, Yucheng Wang +8
We present LiveLight, the first diffusion-based framework for real-time streaming video relighting with interactive 3D lighting control. Achieving this is non-trivial, as it requir…
Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
Jianmin Chen, Jiaqi Tang, Wei Wei +9
Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthen, models may gradually rely…
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Guoxuan Chen, Chufeng Xiao, Haoran Yang +30
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers compet…