209 citations · 919 across the 46 of their papers we have counts for
27 papers · 1 filter
Any4D: Unified Feed-Forward Metric 4D Reconstruction
Jay Karhade, Nikhil Keetha, Yuchen Zhang +4
We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N…
CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives
Zihan Wang, Jiashun Wang, Jeff Tan +4
We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human-scene reconstruction relies on data-driven pr…
PAI-Bench: A Comprehensive Benchmark For Physical AI
Fengzhe Zhou, Jiannan Huang, Jialuo Li +2
Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models…
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
Chancharik Mitra, Yusen Luo, Raj Saravanan +7
Vision-Language Action (VLAs) models promise to extend the remarkable success of vision-language models (VLMs) to robotics. Yet, unlike VLMs in the vision-language domain, VLAs for…
RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
Isaac Robinson, Peter Robicheaux, Matvei Popov +2
Open-vocabulary detectors achieve impressive performance on COCO, but often fail to generalize to real-world datasets with out-of-distribution classes not typically found in their…
HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
Haoyang Weng, Yitang Li, Nikhil Sobanbabu +5
Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for In…