collaborators

10 papers

cs.CV2026

Scene and Human in One World: Reconstruction in a Feedforward Pass

Boao Shi, Qiao Feng, Yiming Huang +1

Reconstructing humans in dynamic scenes from moving monocular cameras remains challenging due to scale ambiguity, human-scene misalignment, and occlusion interference. Rather than…

cs.CV2026

UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching

Qilin Huang, Quynh Anh Huynh, Long Le +5

Existing feed-forward networks excel at predicting a single set of physical properties from visual appearance, but this point-estimate paradigm fundamentally fails to capture the r…

cs.RO2026

PointAction: 3D Points as Universal Action Representations for Robot Control

Mutian Tong, Han Jiang, Qiao Feng +2

Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. How…

cs.RO2026

OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies

Yunzhou Song, Long Le, Yong-Hyun Park +7

Vision-language-action(VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on mo…

cs.CV2025

PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation

Chen Wang, Chuhao Chen, Yiming Huang +4

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limit…

cs.CV2025

Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation

Qingxuan Wu, Zhiyang Dou, Chuan Guo +5

Modeling human-human interactions from text remains challenging because it requires not only realistic individual dynamics but also precise, text-consistent spatiotemporal coupling…