156 citations · 487 across the 39 of their papers we have counts for
38 papers · 1 filter
3DTopia: Large Text-to-3D Generation Model with Hybrid Diffusion Priors
Fangzhou Hong, Jiaxiang Tang, Ziang Cao +8
We present a two-stage text-to-3D generation system, namely 3DTopia, which generates high-quality general 3D assets within 5 minutes using hybrid diffusion priors. The first stage…
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Bo Li, Yuanhan Zhang, Liangyu Chen +5
High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language…
GP-UNIT: Generative Prior for Versatile Unsupervised Image-to-Image Translation
Shuai Yang, Liming Jiang, Ziwei Liu +1
Recent advances in deep learning have witnessed many successful unsupervised image-to-image translation models that learn correspondences between two visual domains without paired…
SAD: Segment Any RGBD
Jun Cen, Yizheng Wu, Kewei Wang +6
The Segment Anything Model (SAM) has demonstrated its effectiveness in segmenting any part of 2D RGB images. However, SAM exhibits a stronger emphasis on texture information while…
RenderMe-360: A Large Digital Asset Library and Benchmarks Towards High-fidelity Head Avatars
Dongwei Pan, Long Zhuo, Jingtan Piao +13
Synthesizing high-fidelity head avatars is a central problem for computer vision and graphics. While head avatar synthesis algorithms have advanced rapidly, the best ones still fac…
ConsistentNeRF: Enhancing Neural Radiance Fields with 3D Consistency for Sparse View Synthesis
Shoukang Hu, Kaichen Zhou, Kaiyu Li +6
Neural Radiance Fields (NeRF) has demonstrated remarkable 3D reconstruction capabilities with dense view images. However, its performance significantly deteriorates under sparse vi…