4 citations · 8 across the 4 of their papers we have counts for
4 papers
Advancing Visual Grounding with Scene Knowledge: Benchmark and Method
Zhihong Chen, Ruifei Zhang, Yibing Song +2
Visual grounding (VG) aims to establish fine-grained alignment between vision and language. Ideally, it can be a testbed for vision-and-language models to evaluate their understand…
Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation
Zunnan Xu, Zhihong Chen, Yong Zhang +3
Parameter Efficient Tuning (PET) has gained attention for reducing the number of parameters while maintaining performance and providing better hardware resource savings, but few st…
On the Difference of BERT-style and CLIP-style Text Encoders
Zhihong Chen, Guiming Hardy Chen, Shizhe Diao +2
Masked language modeling (MLM) has been one of the most popular pretraining recipes in natural language processing, e.g., BERT, one of the representative models. Recently, contrast…
Towards Unifying Medical Vision-and-Language Pre-training via Soft Prompts
Zhihong Chen, Shizhe Diao, Benyou Wang +2
Medical vision-and-language pre-training (Med-VLP) has shown promising improvements on many downstream medical tasks owing to its applicability to extracting generic representation…