activity
20242026
most citedIdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models

4 citations · 4 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CV20264 cited

IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models

Haoxuan You, Rui Sun, Zhecan Wang +5

The field of vision-and-language (VL) understanding has made unprecedented progress with end-to-end large pre-trained VL models (VLMs). However, they still fall short in zero-shot…

cs.CV2026

Compositional Feature Augmentation for Unbiased Scene Graph Generation

Lin Li, Guikun Chen, Jun Xiao +3

Scene Graph Generation (SGG) aims to detect all the visual relation triplets \texttt{sub}, \texttt{pred}, \texttt{obj} in a given image. With the emergence of various advance…

cs.CV2025

Compositional Zero-shot Learning via Progressive Language-based Observations

Lin Li, Guikun Chen, Zhen Wang +2

Compositional zero-shot learning aims to recognize unseen state-object compositions by leveraging known primitives (state and object) during training. However, effectively modeling…

cs.GR2025

View-Consistent 3D Editing with Gaussian Splatting

Yuxuan Wang, Xuanyu Yi, Zike Wu +3

The advent of 3D Gaussian Splatting (3DGS) has revolutionized 3D editing, offering efficient, high-fidelity rendering and enabling precise local manipulations. Currently, diffusion…

cs.CV2024

Decomposed Prototype Learning for Few-Shot Scene Graph Generation

Xingchen Li, Jun Xiao, Guikun Chen +4

Today's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world appli…