65 citations · 100 across the 4 of their papers we have counts for
4 papers
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Tao Gong, Chengqi Lyu, Shilong Zhang +7
We present a vision and language model named MultiModal-GPT to conduct multi-round dialogue with humans. MultiModal-GPT can follow various instructions from humans, such as generat…
Group R-CNN for Weakly Semi-supervised Object Detection with Points
Shilong Zhang, Zhuoran Yu, Liyang Liu +3
We study the problem of weakly semi-supervised object detection with points (WSSOD-P), where the training data is combined by a small set of fully annotated images with bounding bo…
Group Fisher Pruning for Practical Network Compression
Liyang Liu, Shilong Zhang, Zhanghui Kuang +7
Network compression has been widely studied since it is able to reduce the memory and computation cost during inference. However, previous methods seldom deal with complicated stru…
Scale-Equalizing Pyramid Convolution for Object Detection
Xinjiang Wang, Shilong Zhang, Zhuoran Yu +2
Feature pyramid has been an efficient method to extract features at different scales. Development over this method mainly focuses on aggregating contextual information at different…