From the 1 of 19 linked papers with an AI index.
19 papers
ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation
Yu Zhang, Yidi Shao, Wenqi Ouyang +5
ClothTransformer reformulates cloth simulation as an autoregressive sequence modeling problem in a learned latent space, using a unified Transformer architecture that handles diver…
SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration
Jeonghwan Kim, Yushi Lan, Yongwei Chen +3
Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual…
Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models
Armando Fortes, Tianyi Wei, Shangchen Zhou +1
Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textual prompts; however, while traditional…
Boosting Monocular Metric Depth Estimation via Bokeh Rendering
Hangwei Zhang, Armando Fortes, Tianyi Wei +1
Bokeh rendering and depth estimation share a fundamental optical connection, yet existing methods fail to fully exploit this reciprocity. Conventional bokeh pipelines rely heavily…
4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
Yihang Luo, Shangchen Zhou, Yushi Lan +2
We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple motion from geometry or produce lim…
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
Yongwei Chen, Tianyi Wei, Yushi Lan +4
The rapid progress of large multimodal models has inspired efforts toward unified frameworks that couple understanding and generation. While such paradigms have shown remarkable su…