7 papers
DuoGen: Towards General Purpose Interleaved Multimodal Generation
Min Shi, Xiaohui Zeng, Jiannan Huang +13
Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts f…
SpikeVAEDiff: Neural Spike-based Natural Visual Scene Reconstruction via VD-VAE and Versatile Diffusion
Jialu Li, Taiyan Zhou
Reconstructing natural visual scenes from neural activity is a key challenge in neuroscience and computer vision. We present SpikeVAEDiff, a novel two-stage framework that combines…
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
Jitesh Jain, Jialuo Li, Zixian Ma +7
As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this…
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
Jialuo Li, Bin Li, Jiahao Li +1
The application of Large Multimodal Models (LMMs) to long-form video understanding is constrained by limited context lengths and the computationally prohibitive cost of processing…
PAI-Bench: A Comprehensive Benchmark For Physical AI
Fengzhe Zhou, Jiannan Huang, Jialuo Li +2
Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models…
ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling
Qisen Wang, Yifan Zhao, Peisen Shen +2
Although prevailing camera-controlled video generation models can produce cinematic results, lifting them directly to the generation of 3D-consistent and high-fidelity time-synchro…