17 citations · 22 across the 9 of their papers we have counts for
6 papers · 1 filter
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection
Zhuoxiao Chen, Zixin Wang, Yadan Luo +2
LiDAR-based 3D object detection has seen impressive advances in recent times. However, deploying trained 3D detectors in the real world often yields unsatisfactory performance when…
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
Haodong Hong, Sen Wang, Zi Huang +2
Current Vision-and-Language Navigation (VLN) tasks mainly employ textual instructions to guide agents. However, being inherently abstract, the same textual instruction can be assoc…
No Token Left Behind: Efficient Vision Transformer via Dynamic Token Idling
Xuwei Xu, Changlin Li, Yudong Chen +3
Vision Transformers (ViTs) have demonstrated outstanding performance in computer vision tasks, yet their high computational complexity prevents their deployment in computing resour…
Object Detection Difficulty: Suppressing Over-aggregation for Faster and Better Video Object Detection
Bingqing Zhang, Sen Wang, Yifan Liu +3
Current video object detection (VOD) models often encounter issues with over-aggregation due to redundant aggregation strategies, which perform feature aggregation on every frame.…
Cal-SFDA: Source-Free Domain-adaptive Semantic Segmentation with Differentiable Expected Calibration Error
Zixin Wang, Yadan Luo, Zhi Chen +2
The prevalence of domain adaptive semantic segmentation has prompted concerns regarding source domain data leakage, where private information from the source domain could inadverte…
Action2video: Generating Videos of Human 3D Actions
Chuan Guo, Xinxin Zuo, Sen Wang +4
We aim to tackle the interesting yet challenging problem of generating videos of diverse and natural human motions from prescribed action categories. The key issue lies in the abil…