4 citations · 4 across the 7 of their papers we have counts for
6 papers · 1 filter
Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization
Zhu Xu, Jiaqi Tang, Pokai Chen +2
Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justifications. This expanded scope…
Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models
Jiaqi Tang, Shaoyang Zhang, Fandong Zhang +3
The growing number of medical vision foundation models highlights the need for effective model selection. However, mainstream selection methods rely on exhaustive fine-tuning, whic…
Investigating Domain Gaps for Indoor 3D Object Detection
Zijing Zhao, Zhu Xu, Qingchao Chen +2
As a fundamental task for indoor scene understanding, 3D object detection has been extensively studied, and the accuracy on indoor point cloud data has been substantially improved.…
TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge Transferring
Zhu Xu, Ting Lei, Zhimin Li +4
Dynamic Scene Graph Generation (DSGG) aims to create a scene graph for each video frame by detecting objects and predicting their relationships. Weakly Supervised DSGG (WS-DSGG) re…
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
Minghang Zheng, Jiahua Zhang, Qingchao Chen +2
Visual grounding aims to localize the object referred to in an image based on a natural language query. Although progress has been made recently, accurately localizing target objec…
Training-free Video Temporal Grounding using Large-scale Pre-trained Models
Minghang Zheng, Xinhao Cai, Qingchao Chen +2
Video temporal grounding aims to identify video segments within untrimmed videos that are most relevant to a given natural language query. Existing video temporal localization mode…