4 citations · 4 across the 1 of their papers we have counts for
5 papers
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
Haoxuan You, Rui Sun, Zhecan Wang +5
The field of vision-and-language (VL) understanding has made unprecedented progress with end-to-end large pre-trained VL models (VLMs). However, they still fall short in zero-shot…
Compositional Feature Augmentation for Unbiased Scene Graph Generation
Lin Li, Guikun Chen, Jun Xiao +3
Scene Graph Generation (SGG) aims to detect all the visual relation triplets \texttt{sub}, \texttt{pred}, \texttt{obj} in a given image. With the emergence of various advance…
Compositional Zero-shot Learning via Progressive Language-based Observations
Lin Li, Guikun Chen, Zhen Wang +2
Compositional zero-shot learning aims to recognize unseen state-object compositions by leveraging known primitives (state and object) during training. However, effectively modeling…
View-Consistent 3D Editing with Gaussian Splatting
Yuxuan Wang, Xuanyu Yi, Zike Wu +3
The advent of 3D Gaussian Splatting (3DGS) has revolutionized 3D editing, offering efficient, high-fidelity rendering and enabling precise local manipulations. Currently, diffusion…
Decomposed Prototype Learning for Few-Shot Scene Graph Generation
Xingchen Li, Jun Xiao, Guikun Chen +4
Today's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world appli…