17 citations · 35 across the 8 of their papers we have counts for
8 papers · 1 filter
SEAGULL: No-reference Image Quality Assessment for Regions of Interest via Vision-Language Instruction Tuning
Zewen Chen, Juan Wang, Wen Wang +8
Existing Image Quality Assessment (IQA) methods achieve remarkable success in analyzing quality for overall image, but few works explore quality analysis for Regions of Interest (R…
EA-VTR: Event-Aware Video-Text Retrieval
Zongyang Ma, Ziqi Zhang, Yuxin Chen +8
Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often l…
ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual Tracking
Yutong Kou, Jin Gao, Bing Li +4
Recently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter fe…
Component-aware anomaly detection framework for adjustable and logical industrial visual inspection
Tongkun Liu, Bing Li, Xiao Du +4
Industrial visual inspection aims at detecting surface defects in products during the manufacturing process. Although existing anomaly detection models have shown great performance…
PIC 4th Challenge: Semantic-Assisted Multi-Feature Encoding and Multi-Head Decoding for Dense Video Captioning
Yifan Lu, Ziqi Zhang, Yuxin Chen +3
The task of Dense Video Captioning (DVC) aims to generate captions with timestamps for multiple events in one video. Semantic information plays an important role for both localizat…
Cross-Architecture Knowledge Distillation
Yufan Liu, Jiajiong Cao, Bing Li +3
Transformer attracts much attention because of its ability to learn global relations and superior performance. In order to achieve higher performance, it is natural to distill comp…