19 papers
DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
Hanshu Yao, Janfeng Zhong, Niu Lian +1
Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-…
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
Jiarui Yang, Jiale Zhange, Jiawei Li +5
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by…
Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering
Dongxing Mao, Jinpeng Wang, Jiahao Tang +6
Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation abilit…
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
Xingyu Lu, Jinpeng Wang, Yi-Fan Zhang +13
Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for captioning, MLLMs have achieve…
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
Jun Li, Peifeng Lai, Xuhang Lou +5
Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries an…
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Niu Lian, Yuting Wang, Hanshu Yao +5
While multimodal large language models have demonstrated impressive short-term reasoning, they struggle with long-horizon video understanding due to limited context windows and sta…