3 papers
cs.CV2026
From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation
Zhefan Rao, Bin Zou, Xuanhua He +5
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-…
cs.CV2026
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Guoxuan Chen, Chufeng Xiao, Haoran Yang +30
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers compet…
cs.CV2026
InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation
Zhefan Rao, Bin Zou, Haoxuan Che +5
Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same t…