works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.CV2026

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

Guoxuan Chen, Chufeng Xiao, Haoran Yang +30

Boogu-Image-0.1 is an open-source multimodal model family that supports high-quality text-to-image generation, fast inference, instruction-based image editing, and bilingual (Chine…

cs.CV2026

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

Fang Liu, Jinpeng Chen, Ke Xu +7

While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remain…

cs.CV2026

EgoCS-400K: An Egocentric Gameplay Dataset for World Models

Rongjin Guo, Dong Liang, Yuhao Liu +4

The shift from video generation to interactive world modeling places new demands on data: beyond captioned videos, world models require temporally aligned video-action-language tra…

cs.CV2026

World-Shaper: A Unified Framework for 360° Panoramic Editing

Dong Liang, Yuhao Liu, Jinyuan Jia +2

Being able to edit panoramic images is crucial for creating realistic 360° visual experiences. However, existing perspective-based image editing methods fail to model the spatial…

cs.CV2025

Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy

Yuhao Liu, Tengfei Wang, Fang Liu +2

Recent advances in deep generative modeling have unlocked unprecedented opportunities for video synthesis. In real-world applications, however, users often seek tools to faithfully…

cs.CV2025

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

Zaiquan Yang, Yuhao Liu, Gerhard Hancke +1

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large lang…