3 citations · 3 across the 7 of their papers we have counts for
10 papers · 1 filter
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Hohin Kwan, Hongyu Li, Ray Zhang +5
Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events i…
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Houyuan Chen, Hong Li, Xianghao Kong +8
Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often train separate models for each…
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models
Songlin Yang, Xianghao Kong, Anyi Rao
Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, howeve…
InstanceAnimator: Multi-Instance Sketch Video Colorization
Yinhan Zhang, Yue Ma, Bingyuan Wang +5
We propose InstanceAnimator, a novel Diffusion Transformer framework for multi-instance sketch video colorization. Existing methods suffer from three core limitations: inflexible u…
Controllable Text-to-Motion Generation via Modular Body-Part Phase Control
Minyue Dai, Ke Fan, Anyi Rao +2
Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherenc…
Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models
Songchun Zhang, Zeyue Xue, Siming Fu +6
Distilled autoregressive (AR) video models enable efficient streaming generation but frequently misalign with human visual preferences. Existing reinforcement learning (RL) framewo…