1 citations · 1 across the 7 of their papers we have counts for
5 papers · 1 filter
Dive Into the Implicit Biases of Low-rank Vision-language Alignment
Mingjia Shi, Shuo Wang, Xiaobo Wang +7
Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates.…
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Wei Dong, Tianyu Fu, Zhe Yu +9
As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion…
Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly Detection
Lei Fan, Junjie Huang, Donglin Di +4
For anomaly detection (AD), early approaches often train separate models for individual classes, yielding high performance but posing challenges in scalability and resource managem…
Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding
Minghui Wu, Chenxu Zhao, Anyang Su +8
Understanding of video creativity and content often varies among individuals, with differences in focal points and cognitive levels across different ages, experiences, and genders.…
ToCoAD: Two-Stage Contrastive Learning for Industrial Anomaly Detection
Yun Liang, Zhiguang Hu, Junjie Huang +3
Current unsupervised anomaly detection approaches perform well on public datasets but struggle with specific anomaly types due to the domain gap between pre-trained feature extract…