activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

Guoxuan Chen, Chufeng Xiao, Haoran Yang +30

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers compet…

cs.CV2026

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

Fang Liu, Jinpeng Chen, Ke Xu +7

While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remain…

cs.CV2026

EgoCS-400K: An Egocentric Gameplay Dataset for World Models

Rongjin Guo, Dong Liang, Yuhao Liu +4

The shift from video generation to interactive world modeling places new demands on data: beyond captioned videos, world models require temporally aligned video-action-language tra…

cs.CV2026

World-Shaper: A Unified Framework for 360° Panoramic Editing

Dong Liang, Yuhao Liu, Jinyuan Jia +2

Being able to edit panoramic images is crucial for creating realistic 360° visual experiences. However, existing perspective-based image editing methods fail to model the spatial s…

cs.CV2025

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

Zaiquan Yang, Yuhao Liu, Gerhard Hancke +1

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large lang…

cs.CV2025

Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation

Tianyu Huang, Wangguandong Zheng, Tengfei Wang +8

Real-world applications like video gaming and virtual reality often demand the ability to model 3D scenes that users can explore along custom camera trajectories. While significant…