24 papers · 1 filter
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
Shihao Yuan, Yuanze Li, Ruyi Zhang +2
Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue t…
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
Feng Zhu, Shuyang Xie, Yihan Zeng +2
Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by composing multiple task-specifi…
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
Haodong Yu, Yabo Zhang, Donglin Di +2
While diffusion models excel at generating images with conventional dimensions, pushing them to synthesize ultra-high-resolution imagery at extreme aspect ratios (EAR) often trigge…
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
Weijun Zhuang, Yuqing Huang, Weikang Meng +5
Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked vi…
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
Haoyu Wang, Hao Tang, Donglin Di +5
Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, lon…
Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer
Yecong Wan, Mingwen Shao, Renlong Wu +1
In this work, we present Color3D, a highly adaptable framework for colorizing both static and dynamic 3D scenes from monochromatic inputs, delivering visually diverse and chromatic…