From the 1 of 20 linked papers with an AI index.
20 papers
ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation
Yu Zhang, Yidi Shao, Wenqi Ouyang +5
ClothTransformer reformulates cloth simulation as an autoregressive sequence modeling problem in a learned latent space, using a unified Transformer architecture that handles diver…
Syn4D: A Multiview Synthetic 4D Dataset
Zeren Jiang, Yushi Lan, Yihang Luo +8
Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by th…
PhysiFormer: Learning to Simulate Mechanics in World Space
Yiming Chen, Yushi Lan, Andrea Vedaldi
We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models that operate in view-dependent pixel space, PhysiFormer represe…
SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration
Jeonghwan Kim, Yushi Lan, Yongwei Chen +3
Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual…
Direct 3D-Aware Object Insertion via Decomposed Visual Proxies
Jingbo Gong, Yikai Wang, Yushi Lan +6
Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formu…
4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
Yihang Luo, Shangchen Zhou, Yushi Lan +2
We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple motion from geometry or produce lim…