382 citations · 389 across the 13 of their papers we have counts for
9 papers · 1 filter
Thinking with Images via Self-Calling Agent
Wenxi Yang, Yuzhong Zhao, Fang Wan +1
Thinking-with-images paradigms have showcased remarkable visual reasoning capability by integrating visual information as dynamic elements into the Chain-of-Thought (CoT). However,…
DocReward: A Document Reward Model for Structuring and Stylizing
Junpeng Liu, Yuzhong Zhao, Bowen Cao +17
Recent agentic workflows automate professional document generation but focus narrowly on textual quality, overlooking structural and stylistic professionalism, which is equally cri…
Model as a Game: On Numerical and Spatial Consistency for Generative Games
Jingye Chen, Yuzhong Zhao, Yupan Huang +5
Recent advances in generative models have significantly impacted game generation. However, despite producing high-quality graphics and adequately receiving player input, existing m…
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
Feng Liu, Shiwei Zhang, Xiaofeng Wang +6
As a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the mode…
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
Mingxiang Liao, Hannan Lu, Xinyu Zhang +6
Comprehensive and constructive evaluation protocols play an important role in the development of sophisticated text-to-video (T2V) generation models. Existing evaluation protocols…
DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
Yuzhong Zhao, Feng Liu, Yue Liu +4
One fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptabi…