10 papers
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
Zhihao He, Tieyuan Chen, Kangyu Wang +6
Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this…
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
Ziran Qin, Youru Lv, Mingbao Lin +4
Autoregressive (AR) visual generation has emerged as a powerful paradigm for image and multimodal synthesis, owing to its scalability and generality. However, existing AR image gen…
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +5
Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feat…
DND: Boosting Large Language Models with Dynamic Nested Depth
Tieyuan Chen, Xiaodong Chen, Haoxing Chen +3
We introduce Dynamic Nested Depth (DND), a novel method that improves performance for off-the-shelf LLMs by selecting critical tokens to reprocess in a nested depth manner. Specifi…
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
Zhihao He, Tianyao He, Yun Xu +5
Despite the prosperity of the video language model, the current pursuit of comprehensive video reasoning is thwarted by the inherent spatio-temporal incompleteness within individua…
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
Tieyuan Chen, Huabin Liu, Yi Wang +8
Video Question Answering (VideoQA) aims to answer natural language questions based on the given video, with prior work primarily focusing on identifying the duration of relevant se…