activity
20232026
most citedA Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation

8 citations · 15 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2026

Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning

Zhongxiao Cong, Qitao Zhao, Minsik Jeon +1

Current feed-forward 3D/4D reconstruction systems rely on dense geometry and pose supervision -- expensive to obtain at scale and particularly scarce for dynamic real-world scenes.…

cs.CV20251 cited

E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training

Qitao Zhao, Hao Tan, Qianqian Wang +5

Self-supervised pre-training has driven rapid progress in foundation models for language, 2D images, and video, yet remains largely unexplored for learning 3D-aware representations…

cs.CV2025

DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion

Qitao Zhao, Amy Lin, Jeff Tan +3

Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In…

cs.CV2024

Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis

Qitao Zhao, Shubham Tulsiani

Inferring the 3D structure underlying a set of multi-view images typically requires solving two co-dependent tasks -- accurate 3D reconstruction requires precise camera poses, and…

cs.CV20238 cited

A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation

Qitao Zhao, Ce Zheng, Mengyuan Liu +1

The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for…

cs.CV20236 cited

PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation

Qitao Zhao, Ce Zheng, Mengyuan Liu +2

Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relation…