4 papers
Veda: Scalable Video Diffusion via Distilled Sparse Attention
Shihao Han, Hao Yang, Xinting Hu +3
Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse attention methods degrade under…
SemanticNVS: Improving Semantic Scene Understanding in Generative Novel View Synthesis
Xinya Chen, Christopher Wewer, Jiahao Xie +2
We present SemanticNVS, a camera-conditioned multi-view diffusion model for novel view synthesis (NVS), which improves generation quality and consistency by integrating pre-trained…
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
Yifei Yu, Xiaoshan Wu, Xinting Hu +8
Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challe…
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
Bo Wang, Jiehong Lin, Chenzhi Liu +5
We present MG-Nav (Memory-Guided Navigation), a dual-scale framework for zero-shot visual navigation that unifies global memory-guided planning with local geometry-enhanced control…