3 papers
cs.CV2025
Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning
Jiaao Yu, Mingjie Han, Jinkun Jiang +3
The high cost of data annotation has spurred research on training deep learning models in data-limited scenarios. Existing paradigms, however, fail to balance cross-domain transfer…
cs.CV2025
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
Jiaao Yu, Shenwei Li, Mingjie Han +4
Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet…
cs.CV2025
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
Jiaao Yu, Mingjie Han, Tao Gong +2
With the rapid growth of video data, text-video retrieval technology has become increasingly important in numerous application scenarios such as recommendation and search. Early te…