5 papers
POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation
Ruiyan Gong, Meisheng Zhang, Yuxiang Zhao +12
Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical "final-meters" challenge. Existing Vision-Language Navigat…
Orchestrating Spatial Semantics via a Zone-Graph Paradigm for Intricate Indoor Scene Generation
Meisheng Zhang, Shizhao Sun, Yang Zhao +3
Autonomous 3D indoor scene synthesis breaks down in non-convex rooms with tightly coupled spatial constraints. Data-driven generators lack topological priors for long-horizon plann…
SceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RL
Yang Zhao, Shizhao Sun, Meisheng Zhang +3
Current one-pass 3D scene synthesis methods often suffer from spatial hallucinations, such as collisions, due to a lack of deliberative reasoning. To bridge this gap, we introduce…
DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design Generation
Ziyuan Liu, Shizhao Sun, Danqing Huang +5
Graphic design generation demands a delicate balance between high visual fidelity and fine-grained structural editability. However, existing approaches typically bifurcate into eit…
How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction
Zhaolu Kang, Yingjie He, Kehan Jiang +18
Large language models (LLMs) excel at semantic understanding, yet their ability to reconstruct internal structure from scrambled inputs remains underexplored. Sentence-level restor…