858 citations
- University of Chinese Academy of SciencesCN40 papers
- Shanghai Artificial Intelligence LaboratoryCN35 papers
- Chinese Academy of SciencesCN34 papers
- Peking UniversityCN33 papers
- Tsinghua UniversityCN32 papers
- Institute of AutomationCN28 papers
- Shandong Institute of AutomationCN20 papers
- Beihang UniversityCN15 papers
- Chinese University of Hong KongHK10 papers
- University of Hong KongHK10 papers
- Center for Excellence in Brain Science and Intelligence TechnologyCN9 papers
- Fudan UniversityCN9 papers
65 papers · 1 filter
IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning
Jiapeng Li, Ping Wei, Wenjuan Han +2
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark mat…
Revisiting Shape and Texture Reliance with Category-Separability-Calibrated Suppression
Ning Jiang, Tianyi Luo, Zhengyong Huang +1
Feature-suppression evaluations infer model reliance on shape or texture from the accuracy loss caused by attenuating each type of information. Such losses, however, conflate featu…
Miles: Metric Learning with Expandable Subspace for Pre-Trained Model-Based Class-Incremental Learning
Kai Jiang, Zisong Lin, Hongyuan Zhang +2
Class Incremental Learning (CIL) aims to learn new concepts consistently from a data stream without forgetting. Unlike typical CIL methods which need to learn a model from scratch,…
LDFE: Laplacian Decoupled Feature Enhancement Block for Dual-Stream CNN-based RGB-IR Object Detection
Wenhao Dong, Xiaoyan Luo, Linlin Yang +4
The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Existing methods prefer dual-stream CNN bac…
TriDP-PTM: a three-stage distortion-perception tradeoff guides the pre-training model for radar cardiac sensing
Jinye Li, Aidong Men, Yang Liu +1
Cardiovascular diseases (CVDs) remain a leading cause of death globally, necessitating continuous, accurate non-invasive cardiac monitoring. While non-contact radar-based approache…
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Houyuan Chen, Hong Li, Xianghao Kong +8
Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often train separate models for each…