8 papers
Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?
Chengan Che, Chao Wang, Jiayuan Huang +2
Recent advancements in self-supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extending these visual foundations t…
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
Chao Wang, Xudong Tan, Jianjian Cao +2
Multimodal Large Language Models have achieved significant success in offline video understanding, yet their application to streaming videos is severely limited by the linear explo…
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
Bo Yu, Fengze Yang, Yiming Liu +6
The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven infe…
FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models
Haoyang Li, Liang Wang, Siyu Zhou +5
CLIP-based prompt tuning enables pretrained Vision-Language Models (VLMs) to efficiently adapt to downstream tasks. Although existing studies have made significant progress, they p…
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework
Chao Wang, Chunbai Zhang, Yongxiao Tian +2
Visual reasoning refers to the task of solving questions about visual information. Current visual reasoning methods typically employ pre-trained vision-language model (VLM) strateg…
TPC: Cross-Temporal Prediction Connection for Vision-Language Model Hallucination Reduction
Chao Wang, Weiwei Fu, Yang Zhou
Vision-language models (VLMs) have achieved remarkable advancements, capitalizing on the impressive capabilities of large language models (LLMs) across diverse tasks. Despite this,…