most citedInstant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model

32 citations · 64 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20237 cited

PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction

Peng Wang, Hao Tan, Sai Bi +6

We propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating…

cs.CV202332 cited

Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model

Jiahao Li, Hao Tan, Kai Zhang +7

Text-to-3D with diffusion models has achieved remarkable progress in recent years. However, existing methods either rely on score distillation-based optimization which suffer from…

cs.CV202318 cited

DMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction Model

Yinghao Xu, Hao Tan, Fujun Luan +8

We propose \textbf{DMV3D}, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model inco…

cs.LG20235 cited

Federated Skewed Label Learning with Logits Fusion

Yuwei Wang, Runhan Li, Hao Tan +5

Federated learning (FL) aims to collaboratively train a shared model across multiple clients without transmitting their local data. Data heterogeneity is a critical challenge in re…

cs.SD20222 cited

Adversarial Attacks on ASR Systems: An Overview

Xiao Zhang, Hao Tan, Xuan Huang +3

With the development of hardware and algorithms, ASR(Automatic Speech Recognition) systems evolve a lot. As The models get simpler, the difficulty of development and deployment bec…

cs.CV2022

CLEAR: Improving Vision-Language Navigation with Cross-Lingual, Environment-Agnostic Representations

Jialu Li, Hao Tan, Mohit Bansal

Vision-and-Language Navigation (VLN) tasks require an agent to navigate through the environment based on language instructions. In this paper, we aim to solve two key challenges in…