activity
20222025
most citedFERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos

6 citations · 9 across the 12 of their papers we have counts for

collaborators

15 papers

cs.CV2025

RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations

Xingqi He, Yujie Zhang, Shuyong Gao +6

Text-guided object segmentation requires both cross-modal reasoning and pixel grounding abilities. Most recent methods treat text-guided segmentation as one-shot grounding, where t…

cs.CV2025

Collaborative Reconstruction and Repair for Multi-class Industrial Anomaly Detection

Qishan Wang, Haofeng Wang, Shuyong Gao +5

Industrial anomaly detection is a challenging open-set task that aims to identify unknown anomalous patterns deviating from normal data distribution. To avoid the significant memor…

cs.CV2025

Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory

Yuxuan Lin, Hanjing Yan, Xuan Tong +6

Few-shot multimodal industrial anomaly detection is a critical yet underexplored task, offering the ability to quickly adapt to complex industrial scenarios. In few-shot settings,…

cs.CV2025

Search is All You Need for Few-shot Anomaly Detection

Qishan Wang, Jia Guo, Shuyong Gao +5

Few-shot anomaly detection (FSAD) has emerged as a crucial yet challenging task in industrial inspection, where normal distribution modeling must be accomplished with only a few no…

cs.CV2025

HSS-IAD: A Heterogeneous Same-Sort Industrial Anomaly Detection Dataset

Qishan Wang, Shuyong Gao, Junjie Hu +4

Multi-class Unsupervised Anomaly Detection algorithms (MUAD) are receiving increasing attention due to their relatively low deployment costs and improved training efficiency. Howev…

cs.CV2025

Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos

Yuang Feng, Shuyong Gao, Fuzhen Yan +4

Video Camouflaged Object Detection (VCOD) aims to segment objects whose appearances closely resemble their surroundings, posing a challenging and emerging task. Existing vision mod…