289 citations · 636 across the 34 of their papers we have counts for
32 papers · 1 filter
Efficient Test-Time Prompt Tuning for Vision-Language Models
Yuhan Zhu, Guozhen Zhang, Chen Xu +4
Vision-language models have showcased impressive zero-shot classification capabilities when equipped with suitable text prompts. Previous studies have shown the effectiveness of te…
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Qingyun Li, Zhe Chen, Weiyun Wang +37
Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resem…
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
Xinhao Li, Zhenpeng Huang, Jing Wang +2
With the growth of high-quality data and advancement in visual pre-training paradigms, Video Foundation Models (VFMs) have made significant progress recently, demonstrating their r…
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
Yuhan Zhu, Yuyang Ji, Zhiyu Zhao +2
Pre-trained vision-language models (VLMs) have shown impressive results in various visual classification tasks. However, we often fail to fully unleash their potential when adaptin…
STMixer: A One-Stage Sparse Action Detector
Tao Wu, Mengqi Cao, Ziteng Gao +2
Traditional video action detectors typically adopt the two-stage pipeline, where a person detector is first employed to generate actor boxes and then 3D RoIAlign is used to extract…
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
Tao Wu, Runyu He, Gangshan Wu +1
Video-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation…