1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
Zhipeng Huang, Zhizheng Zhang, Zheng-Jun Zha +2
The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very…
Vector-based Representation is the Key: A Study on Disentanglement and Compositional Generalization
Tao Yang, Yuwang Wang, Cuiling Lan +2
Recognizing elementary underlying concepts from observations (disentanglement) and generating novel combinations of these concepts (compositional generalization) are fundamental ab…
Structural Multiplane Image: Bridging Neural View Synthesis and 3D Reconstruction
Mingfang Zhang, Jinglu Wang, Xiao Li +3
The Multiplane Image (MPI), containing a set of fronto-parallel RGBA layers, is an effective and efficient representation for view synthesis from sparse inputs. Yet, its fixed stru…
Unifying Layout Generation with a Decoupled Diffusion Model
Mude Hui, Zhizheng Zhang, Xiaoyi Zhang +3
Layout generation aims to synthesize realistic graphic scenes consisting of elements with different attributes including category, size, position, and between-element relation. It…