5 papers
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
Zhicheng Zhang, Wentao Gu, Weicheng Wang +5
Omnimodal understanding entails a massive, highly redundant search space of cross-modal interactions, demanding focused and deliberative reasoning. Current reasoning paradigms rely…
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
Chenxi Zhao, Chen Zhu, Xiaokun Feng +6
Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existing research on MeanFlow prima…
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
Weicheng Wang, Zhicheng Zhang, Zhongqi Zhang +6
Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructing paired video editing data…
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
Zhicheng Zhang, Weicheng Wang, Yongjie Zhu +4
Understanding and predicting emotion from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While adva…
MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
Zhicheng Zhang, Wuyou Xia, Chenxi Zhao +7
Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by a generalizable attention architecture. Advance…