14 citations · 22 across the 4 of their papers we have counts for
6 papers
Power-LLaVA: Large Language and Vision Assistant for Power Transmission Line Inspection
Jiahao Wang, Mingxuan Li, Haichen Luo +4
The inspection of power transmission line has achieved notable achievements in the past few years, primarily due to the integration of deep learning technology. However, current in…
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
Chenyu Yang, Xizhou Zhu, Jinguo Zhu +9
Recently, vision model pre-training has evolved from relying on manually annotated datasets to leveraging large-scale, web-crawled image-text data. Despite these advances, there is…
Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
Hao Li, Jinguo Zhu, Xiaohu Jiang +8
Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to elimi…
Layerwise Optimization by Gradient Decomposition for Continual Learning
Shixiang Tang, Dapeng Chen, Jinguo Zhu +2
Deep neural networks achieve state-of-the-art and sometimes super-human performance across various domains. However, when learning tasks sequentially, the networks easily forget th…
Multiple Domain Experts Collaborative Learning: Multi-Source Domain Generalization For Person Re-Identification
Shijie Yu, Feng Zhu, Dapeng Chen +5
Recent years have witnessed significant progress in person re-identification (ReID). However, current ReID approaches still suffer from considerable performance degradation when un…
Complementary Relation Contrastive Distillation
Jinguo Zhu, Shixiang Tang, Dapeng Chen +5
Knowledge distillation aims to transfer representation ability from a teacher model to a student model. Previous approaches focus on either individual representation distillation o…