7 papers
Video-Browser: Towards Agentic Open-web Video Browsing
Zhengyang Liang, Yan Shu, Xiangrui Liu +5
The evolution of autonomous agents is redefining information seeking, transitioning from passive retrieval to proactive, open-ended web research. However, a significant modality ga…
Reverse Personalization
Han-Wei Kung, Tuomas Varanka, Nicu Sebe
Recent text-to-image diffusion models have demonstrated remarkable generation of realistic facial images conditioned on textual prompts and human identities, enabling creating pers…
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
Ji Zhang, Shihan Wu, Lianli Gao +3
Despite the great promise of Prompt Tuning (PT) in adapting large Vision-Language Pretrained Models (VLPMs) to downstream tasks, they often struggle to overcome the Base-New Tradeo…
Reliable Few-shot Learning under Dual Noises
Ji Zhang, Jingkuan Song, Lianli Gao +2
Recent advances in model pre-training give rise to task adaptation-based few-shot learning (FSL), where the goal is to adapt a pre-trained task-agnostic model for capturing task-sp…
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
Yan Shu, Hangui Lin, Yexin Liu +7
Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, th…
VidText: Towards Comprehensive Evaluation for Video Text Understanding
Zhoufaran Yang, Yan Shu, Jing Wang +8
Visual texts embedded in videos carry rich semantic information, which is crucial for both holistic video understanding and fine-grained reasoning about local human actions. Howeve…