4 citations · 4 across the 1 of their papers we have counts for
7 papers · 1 filter
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
Haoxuan You, Rui Sun, Zhecan Wang +5
The field of vision-and-language (VL) understanding has made unprecedented progress with end-to-end large pre-trained VL models (VLMs). However, they still fall short in zero-shot…
Compositional Feature Augmentation for Unbiased Scene Graph Generation
Lin Li, Guikun Chen, Jun Xiao +3
Scene Graph Generation (SGG) aims to detect all the visual relation triplets \texttt{sub}, \texttt{pred}, \texttt{obj} in a given image. With the emergence of various advance…
Compositional Zero-shot Learning via Progressive Language-based Observations
Lin Li, Guikun Chen, Zhen Wang +2
Compositional zero-shot learning aims to recognize unseen state-object compositions by leveraging known primitives (state and object) during training. However, effectively modeling…
Decomposed Prototype Learning for Few-Shot Scene Graph Generation
Xingchen Li, Jun Xiao, Guikun Chen +4
Today's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world appli…
Cross-Modal Conditioned Reconstruction for Language-guided Medical Image Segmentation
Xiaoshuang Huang, Hongxiang Li, Meng Cao +3
Recent developments underscore the potential of textual information in enhancing learning models for a deeper understanding of medical visual semantics. However, language-guided me…
A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future
Chaoyang Zhu, Long Chen
As the most fundamental scene understanding tasks, object detection and segmentation have made tremendous progress in deep learning era. Due to the expensive manual labeling cost,…