5 papers
CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents
Qianru Li, Xuyang Chen, Erkin Türköz +5
CinemaTraj generates cinematic camera movements in 3D environments from natural language prompts by using an LLM agent that reasons over a structured 3D scene graph and produces co…
FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning
Ping Zhong, Manling Teng, Tao Wu +3
Mobile manipulation requires robots to identify Floor Affordance (FloAff) that maximizes downstream manipulation success rather than merely ensuring navigation feasibility. FloAff…
Spatially Generalizable Mobile Manipulation via Adaptive Experience Selection and Dynamic Imagination
Ping Zhong, Liangbai Liu, Bolei Chen +4
Mobile Manipulation (MM) involves long-horizon decision-making over multi-stage compositions of heterogeneous skills, such as navigation and picking up objects. Despite recent prog…
GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
Antonio Ruiz, Tao Wu, Andrew Melnik +6
Methods that synthesize indoor 3D scenes from text prompts have wide-ranging applications in film production, interior design, video games, virtual reality, and synthetic data gene…
LADB: Latent Aligned Diffusion Bridges for Semi-Supervised Domain Translation
Xuqin Wang, Tao Wu, Yanfeng Zhang +6
Diffusion models excel at generating high-quality outputs but face challenges in data-scarce domains, where exhaustive retraining or costly paired data are often required. To addre…