18 papers
Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation
Mining Tan, Yinuo Wang, Ziqi Zhou +6
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances…
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Ziqi Zhou, Weize Quan, Mining Tan +6
Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grain…
GraphPO: Graph-based Policy Optimization for Reasoning Models
Yuliang Zhan, Xinyu Tang, Jian Li +7
Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses indepe…
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
Shuai Tan, Biao Gong, Yujie Wei +8
Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arbitrary subjects transfer through…
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
Longteng Jiang, DanDan Zheng, Qianqian Qiao +7
The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics…
TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion Generation
Yiyang Cao, Yunze Deng, Ziyu Lin +5
Text-to-motion generation, a rapidly evolving field in computer vision, aims to produce realistic and text-aligned motion sequences. Current methods primarily focus on spatial-temp…