3 papers
cs.CV2024
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
Longfei Huang, Feng Yu, Zhihao Guan +2
This report presents a solution for the zero-shot referring expression comprehension task. Visual-language multimodal base models (such as CLIP, SAM) have gained significant attent…
cs.CV2024
The Solution for Language-Enhanced Image New Category Discovery
Haonan Xu, Dian Chao, Xiangyu Wu +2
Treating texts as images, combining prompts with textual labels for prompt tuning, and leveraging the alignment properties of CLIP have been successfully applied in zero-shot multi…
cs.LG2024
The Solution for the AIGC Inference Performance Optimization Competition
Sishun Pan, Haonan Xu, Zhonghua Wan +1
In recent years, the rapid advancement of large-scale pre-trained language models based on transformer architectures has revolutionized natural language processing tasks. Among the…