60 citations · 92 across the 31 of their papers we have counts for
24 papers · 1 filter
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
Jiarui Yang, Jiale Zhange, Jiawei Li +5
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by…
Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering
Dongxing Mao, Jinpeng Wang, Jiahao Tang +6
Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation abilit…
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
Xingyu Lu, Jinpeng Wang, Yi-Fan Zhang +13
Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for captioning, MLLMs have achieve…
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
Jun Li, Peifeng Lai, Xuhang Lou +5
Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries an…
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Niu Lian, Yuting Wang, Hanshu Yao +5
While multimodal large language models have demonstrated impressive short-term reasoning, they struggle with long-horizon video understanding due to limited context windows and sta…
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
Tianci Luo, Haohao Pan, Jinpeng Wang +5
Visual in-context learning (VICL) enables visual foundation models to handle multiple tasks by steering them with demonstrative prompts. The choice of such prompts largely influenc…