4 papers
LoopVAE: Recurrent Depth Across Scales for Visual Tokenization
Zhiying Lu
Hierarchical visual tokenizers typically allocate different processing blocks to different spatial scales. We ask how much of this computation can use the same parameters. LoopVAE…
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
Bowei Pu, Chuanbin Liu, Yifan Ge +5
Video Large Language Models improve reasoning over complex videos by generating intermediate reasoning text. However, reliable reasoning depends on accurate video perception. In ex…
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
Yinglu Li, Zhiying Lu, Zhihang Liu +3
Multi-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the en…
From Evaluation to Defense: Advancing Safety in Video Large Language Models
Yiwei Sun, Peiqi Jiang, Chuanbin Liu +3
While the safety risks of image-based large language models (Image LLMs) have been extensively studied, their video-based counterparts (Video LLMs) remain critically under-examined…