3 citations · 5 across the 6 of their papers we have counts for
6 papers · 1 filter
See the Text: From Tokenization to Visual Reading
Ling Xing, Rui Yan, Alex Jinpeng Wang +2
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle ty…
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
Shiyu Xuan, Zechao Li, Jinhui Tang
Multi-modal object tracking integrates auxiliary modalities such as depth, thermal infrared, event flow, and language to provide additional information beyond RGB images, showing g…
A Recover-then-Discriminate Framework for Robust Anomaly Detection
Peng Xing, Dong Zhang, Jinhui Tang +1
Anomaly detection (AD) has been extensively studied and applied in a wide range of scenarios in the recent past. However, there are still gaps between achieved and desirable levels…
Spatial Structure Constraints for Weakly Supervised Semantic Segmentation
Tao Chen, Yazhou Yao, Xingguo Huang +3
The image-level label has prevailed in weakly supervised semantic segmentation tasks due to its easy availability. Since image-level labels can only indicate the existence or absen…
MNet: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action Recognition
Hao Tang, Jun Liu, Shuanglin Yan +3
Due to the scarcity of manually annotated data required for fine-grained video understanding, few-shot fine-grained (FS-FG) action recognition has gained significant attention, wit…
ADPS: Asymmetric Distillation Post-Segmentation for Image Anomaly Detection
Peng Xing, Hao Tang, Jinhui Tang +1
Knowledge Distillation-based Anomaly Detection (KDAD) methods rely on the teacher-student paradigm to detect and segment anomalous regions by contrasting the unique features extrac…