activity
20212026
most citedMISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for Recommendation

60 citations · 92 across the 31 of their papers we have counts for

collaborators
Showing cs.CVShow all

24 papers · 1 filter

cs.CV2026

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

Jiarui Yang, Jiale Zhange, Jiawei Li +5

World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by…

cs.CV2026

Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering

Dongxing Mao, Jinpeng Wang, Jiahao Tang +6

Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation abilit…

cs.CV2026

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

Xingyu Lu, Jinpeng Wang, Yi-Fan Zhang +13

Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for captioning, MLLMs have achieve…

cs.CV2026

Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval

Jun Li, Peifeng Lai, Xuhang Lou +5

Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries an…

cs.CV2026

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Niu Lian, Yuting Wang, Hanshu Yao +5

While multimodal large language models have demonstrated impressive short-term reasoning, they struggle with long-horizon video understanding due to limited context windows and sta…

cs.CV2026

Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning

Tianci Luo, Haohao Pan, Jinpeng Wang +5

Visual in-context learning (VICL) enables visual foundation models to handle multiple tasks by steering them with demonstrative prompts. The choice of such prompts largely influenc…