9 papers · 1 filter
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
Qingyang Liu, Bingjie Gao, Canmiao Fu +9
Recent unified models integrate multimodal understanding and generation within a single framework. However, an "understanding-generation gap" persists, where models can capture use…
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
Bingjie Gao, Qianli Ma, Xiaoxue Wu +9
Prompt design plays a crucial role in text-to-video (T2V) generation, yet user-provided prompts are often short, unstructured, and misaligned with training data, limiting the gener…
Reflection Generation for Composite Image Using Diffusion Model
Haonan Zhao, Qingyang Liu, Jiaxuan Chen +1
Image composition involves inserting a foreground object into the background while synthesizing environment-consistent effects such as shadows and reflections. Although shadow gene…
AnimateScene: Camera-controllable Animation in Any Scene
Qingyang Liu, Bingjie Gao, Weiheng Huang +10
Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plaus…
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
Harold Haodong Chen, Disen Lan, Wen-Jie Shu +10
The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consi…
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
Zelin Peng, Zhengqin Xu, Qingyang Liu +2
Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computation…