9 papers · 1 filter
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
Hangui Lin, Yan Shu, Zhengyang Liang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a leading paradigm for enhancing visual reasoning in Multimodal Large Language Models (MLLMs). However, existin…
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
Jinlong Li, Liyuan Jiang, Haonan Zhang +1
Video Large Language Models (VLLMs) demonstrate strong video understanding but suffer from inefficiency due to redundant visual tokens. Existing pruning primary targets intra-frame…
Video-Browser: Towards Agentic Open-web Video Browsing
Zhengyang Liang, Yan Shu, Xiangrui Liu +5
The evolution of autonomous agents is redefining information seeking, transitioning from passive retrieval to proactive, open-ended web research. However, a significant modality ga…
Reverse Personalization
Han-Wei Kung, Tuomas Varanka, Nicu Sebe
Recent text-to-image diffusion models have demonstrated remarkable generation of realistic facial images conditioned on textual prompts and human identities, enabling creating pers…
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
Ji Zhang, Shihan Wu, Lianli Gao +3
Despite the great promise of Prompt Tuning (PT) in adapting large Vision-Language Pretrained Models (VLPMs) to downstream tasks, they often struggle to overcome the Base-New Tradeo…
Reliable Few-shot Learning under Dual Noises
Ji Zhang, Jingkuan Song, Lianli Gao +2
Recent advances in model pre-training give rise to task adaptation-based few-shot learning (FSL), where the goal is to adapt a pre-trained task-agnostic model for capturing task-sp…