10 citations · 11 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 10 cited
VL-Mamba: Exploring State Space Models for Multimodal Learning
Yanyuan Qiao, Zheng Yu, Longteng Guo +5
Multimodal large language models (MLLMs) have attracted widespread interest and have rich applications. However, the inherent attention mechanism in its Transformer structure requi…
cs.CV2023
GLOBER: Coherent Non-autoregressive Video Generation via GLOBal Guided Video DecodER
Mingzhen Sun, Weining Wang, Zihan Qin +3
Video generation necessitates both global coherence and local realism. This work presents a novel non-autoregressive method GLOBER, which first generates global features to obtain…
cs.CV2023★ 1 cited
MOSO: Decomposing MOtion, Scene and Object for Video Prediction
Mingzhen Sun, Weining Wang, Xinxin Zhu +1
Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their d…