activity
20242026
collaborators

10 papers

cs.CV2026

OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder

Sensen Gao, Zhaoqing Wang, Qihang Cao +5

Existing diffusion-based 3D scene generation methods primarily operate in 2D image/video latent spaces, which makes maintaining cross-view appearance and geometric consistency inhe…

cs.CL2025

IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models

Shaokun Zhang, Xiaobo Xia, Zhaoqing Wang +4

In-context learning is a promising paradigm that utilizes in-context examples as prompts for the predictions of large language models. These prompts are crucial for achieving stron…

cs.CV2025

MF-VITON: High-Fidelity Mask-Free Virtual Try-On with Minimal Input

Zhenchen Wan, Yanwu xu, Dongting Hu +6

Recent advancements in Virtual Try-On (VITON) have significantly improved image realism and garment detail preservation, driven by powerful text-to-image (T2I) diffusion models. Ho…

cs.CV2025

TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On

Zhenchen Wan, Yanwu Xu, Zhaoqing Wang +3

Recent advancements in Virtual Try-On (VTO) have demonstrated exceptional efficacy in generating realistic images and preserving garment details, largely attributed to the robust g…

cs.CV2025

LaVin-DiT: Large Vision Diffusion Transformer

Zhaoqing Wang, Xiaobo Xia, Runnan Chen +4

This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative fra…

cs.CV2024

PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM

Runnan Chen, Zhaoqing Wang, Jiepeng Wang +4

Understanding geometric, semantic, and instance information in 3D scenes from sequential video data is essential for applications in robotics and augmented reality. However, existi…