7 citations · 27 across the 10 of their papers we have counts for
8 papers
RepParser: End-to-End Multiple Human Parsing with Representative Parts
Xiaojia Chen, Xuanhan Wang, Lianli Gao +1
Existing methods of multiple human parsing usually adopt a two-stage strategy (typically top-down and bottom-up), which suffers from either strong dependence on prior detection or…
Towards Open-vocabulary Scene Graph Generation with Prompt-based Finetuning
Tao He, Lianli Gao, Jingkuan Song +1
Scene graph generation (SGG) is a fundamental task aimed at detecting visual relations between objects in an image. The prevailing SGG methods require all object classes to be give…
Prompting for Multi-Modal Tracking
Jinyu Yang, Zhe Li, Feng Zheng +2
Multi-modal tracking gains attention due to its ability to be more accurate and robust in complex scenarios compared to traditional RGB-based tracking. Its key lies in how to fuse…
Frequency Domain Model Augmentation for Adversarial Attack
Yuyang Long, Qilong Zhang, Boheng Zeng +4
For black-box attacks, the gap between the substitute model and the victim model is usually large, which manifests as a weak attack performance. Motivated by the observation that t…
Adaptive Fine-Grained Predicates Learning for Scene Graph Generation
Xinyu Lyu, Lianli Gao, Pengpeng Zeng +2
The performance of current Scene Graph Generation (SGG) models is severely hampered by hard-to-distinguish predicates, e.g., woman-on/standing on/walking on-beach. As general SGG m…
Skeleton-based Action Recognition via Adaptive Cross-Form Learning
Xuanhan Wang, Yan Dai, Lianli Gao +1
Skeleton-based action recognition aims to project skeleton sequences to action categories, where skeleton sequences are derived from multiple forms of pre-detected points. Compared…