activity
20222024
most citedTripoSR: Fast 3D Object Reconstruction from a Single Image

16 citations · 40 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CV2024

Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Le Zhuo, Ruoyi Du, Han Xiao +19

Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers that establishes a unified framework for transforming noise into various modalities, such as images and vi…

cs.LG2024

Exploring Text-to-Motion Generation with Human Preference

Jenny Sheng, Matthieu Lin, Andrew Zhao +5

This paper presents an exploration of preference learning in text-to-motion generation. We find that current improvements in text-to-motion generation still rely on datasets requir…

cs.CV202416 cited

TripoSR: Fast 3D Object Reconstruction from a Single Image

Dmitry Tochilkin, David Pankratz, Zexiang Liu +7

This technical report introduces TripoSR, a 3D reconstruction model leveraging transformer architecture for fast feed-forward 3D generation, producing 3D mesh from a single image i…

cs.CV20238 cited

Text-to-3D with Classifier Score Distillation

Xin Yu, Yuan-Chen Guo, Yangguang Li +3

Text-to-3D generation has made remarkable progress recently, particularly with methods based on Score Distillation Sampling (SDS) that leverages pre-trained 2D diffusion models. Wh…

cs.CV2023

Mask Hierarchical Features For Self-Supervised Learning

Fenggang Liu, Yangguang Li, Feng Liang +3

This paper shows that Masking the Deep hierarchical features is an efficient self-supervised method, denoted as MaskDeep. MaskDeep treats each patch in the representation space as…

cs.CV20237 cited

Fast-BEV: Towards Real-time On-vehicle Bird's-Eye View Perception

Bin Huang, Yangguang Li, Enze Xie +7

Recently, the pure camera-based Bird's-Eye-View (BEV) perception removes expensive Lidar sensors, making it a feasible solution for economical autonomous driving. However, most exi…