3 papers
cs.CV2024
World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
Jiacong Wang, Bohong Wu, Haiyong Jiang +4
Recent advances in Vision-Language Models (VLMs) and the scarcity of high-quality multi-modal alignment data have inspired numerous researches on synthetic VLM data generation. The…
cs.CV2022
PartCom: Part Composition Learning for 3D Open-Set Recognition
Weng Tingyu, Xiao Jun, Jiang Haiyong
3D recognition is the foundation of 3D deep learning in many emerging fields, such as autonomous driving and robotics.Existing 3D methods mainly focus on the recognition of a fixed…
cs.CV2021
CSG-Stump: A Learning Friendly CSG-Like Representation for Interpretable Shape Parsing
Daxuan Ren, Jianmin Zheng, Jianfei Cai +8
Generating an interpretable and compact representation of 3D shapes from point clouds is an important and challenging problem. This paper presents CSG-Stump Net, an unsupervised en…