7 papers · 1 filter
SupScene: Scene-Structured Overlap Supervision for Image Retrieval in Unconstrained SfM
Xulei Shi, Maoyu Wang, Yuning Peng +5
Image retrieval is a critical step for reducing the quadratic cost of image matching in unconstrained Structure-from-Motion (SfM). Unlike generic image retrieval, however, the rele…
TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
Qihang Zhou, Binbin Gao, Guansong Pang +3
Adapting CLIP for anomaly detection on unseen objects has shown strong potential in a zero-shot manner. However, existing methods typically rely on a single textual space to align…
Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment
Xin Wang, Peng-Jie Li, Yuan-Yuan Shen
Long-term action quality assessment (AQA) focuses on evaluating the quality of human activities in videos lasting up to several minutes. This task plays an important role in the au…
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
Liqiang Jing, Guiming Hardy Chen, Ehsan Aghazadeh +2
Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in multimodal tasks, but visual object hallucination remains a persistent issue. It refers to scenarios whe…
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
Jing Gu, Yuwei Fang, Ivan Skorokhodov +4
Video editing serves as a fundamental pillar of digital media, spanning applications in entertainment, education, and professional communication. However, previous methods often ov…
When Do We Not Need Larger Vision Models?
Baifeng Shi, Ziyang Wu, Maolin Mao +2
Scaling up the size of vision models has been the de facto standard to obtain more powerful visual representations. In this work, we discuss the point beyond which larger vision mo…