1 citations · 1 across the 5 of their papers we have counts for
9 papers · 1 filter
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
Jialong Zuo, Haotong Zuo, Shiwei Zhang +5
Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual nar…
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing
Feng Wang, Canmiao Fu, Zhipeng Huang +3
Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all histor…
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
Yiwei Guo, Shaobin Zhuang, Zhipeng Huang +4
Building a unified visual tokenizer is essential for bridging the gap between visual understanding and generation. Yet existing approaches struggle with the inherent conflict betwe…
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
Shaobin Zhuang, Yiwei Guo, Canmiao Fu +7
Visual tokenizer is a critical component for vision generation. However, the existing tokenizers often face unsatisfactory trade-off between compression ratios and reconstruction f…
Text-guided Visual Prompt DINO for Generic Segmentation
Yuchen Guan, Chong Sun, Canmiao Fu +3
Recent advancements in multimodal vision models have highlighted limitations in late-stage feature fusion and suboptimal query selection for hybrid prompts open-world segmentation,…
Video-GPT via Next Clip Diffusion
Shaobin Zhuang, Zhipeng Huang, Ying Zhang +6
GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alte…