21 papers
WorldClaw: Agentic 3D Open-World Generation at Scale
Chunchao Guo, Jinpeng Li, Yang Li +1
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, an…
SceneActBench: Can Agents Act on the 3D Scenes They See?
Yifei Zhao, Xiangxin Zhou, Wenhao Yang +11
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operat…
R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow
Zijie Wu, Lixin Xu, Puhua Jiang +3
Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical deployment faces a critical y…
ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation
Hanxiao Sun, Mingxin Yang, Shuhui Yang +5
Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unseen regions. Multi-view condi…
Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence
Zebin He, Mingxin Yang, Shuhui Yang +4
Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by compressing 3D shape into a glo…
MeshFIM: Local Low-Poly Mesh Editing via Fill-in-the-Middle Autoregressive Generation
Dingdong Yang, Jian Liu, Biwen Lei +6
Autoregressive (AR) models can generate high-quality low-poly meshes from point clouds, but they still operate in an all-or-nothing manner: when a local region is unsatisfactory, t…