4 papers
Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection
Manwen Yang, Leqian Ding, Yu Guo +1
Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to object-centric bias, normal and anomalous text prototypes exhibit a high s…
TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models
Leqian Ding, Junning Qiu, Manwen Yang +2
Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progres…
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
Chenyu Hui, Xiaodi Huang, Siyu Xu +5
Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often s…
One-step Multi-view Clustering With Adaptive Low-rank Anchor-graph Learning
Zhiyuan Xue, Ben Yang, Xuetao Zhang +2
In light of their capability to capture structural information while reducing computing complexity, anchor graph-based multi-view clustering (AGMC) methods have attracted considera…