7 papers
DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement
Kai Wang, Ziheng Ouyang, Xuying Zhang +2
With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches…
GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation
Yuhao Wan, Lijuan Liu, Jingzhi Zhou +6
Previous works that leverage video models for image-to-3D scene generation often suffer from geometric distortions and blurry content. Using video generation models to implicitly m…
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
Bo-Wen Yin, Jiao-Long Cao, Xuying Zhang +3
Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for…
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang +6
Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views…
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
Xuying Zhang, Yutong Liu, Yangguang Li +8
We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-qu…
Referring Camouflaged Object Detection
Xuying Zhang, Bowen Yin, Zheng Lin +3
We consider the problem of referring camouflaged object detection (Ref-COD), a new task that aims to segment specified camouflaged objects based on a small set of referring images…