collaborators

7 papers

cs.CV2026

DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement

Kai Wang, Ziheng Ouyang, Xuying Zhang +2

With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches…

cs.CV2026

GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

Yuhao Wan, Lijuan Liu, Jingzhi Zhou +6

Previous works that leverage video models for image-to-3D scene generation often suffer from geometric distortions and blurry content. Using video generation models to implicitly m…

cs.CV2025

OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

Bo-Wen Yin, Jiao-Long Cao, Xuying Zhang +3

Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for…

cs.CV2025

AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction

Xuying Zhang, Yupeng Zhou, Kai Wang +6

Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views…

cs.CV2025

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

Xuying Zhang, Yutong Liu, Yangguang Li +8

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-qu…

cs.CV2025

Referring Camouflaged Object Detection

Xuying Zhang, Bowen Yin, Zheng Lin +3

We consider the problem of referring camouflaged object detection (Ref-COD), a new task that aims to segment specified camouflaged objects based on a small set of referring images…