9 papers
Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition
Hongyu Qu, Ling Xing, Jiachao Zhang +3
Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level representations for each video by desi…
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
Xun Jiang, Yufan Gu, Disen Hu +5
Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy corruption. While these issues…
Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
Bo Zhou, Qiuxia Lai, Zeren Sun +3
Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such…
Beyond Quadratic: Linear-Time Change Detection with RWKV
Zhenyu Yang, Gensheng Pei, Tao Chen +4
Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependenci…
OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
Hongyu Qu, Jianan Wei, Xiangbo Shu +3
Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to i) the scarcity of annotated datasets, and ii) the insufficient diversity of…
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
Gensheng Pei, Tao Chen, Yujia Wang +4
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot…