19 papers
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Yue Zhang, Yingzhao Jian, Yunqiu Xu +2
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric c…
Large language model agents accelerate inverse design of metal-organic frameworks for gas separation
Zhaolin Hu, Hehe Fan, Wangyihan Guo +4
Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes inverse design difficult under simultaneo…
Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
Yingzhao Jian, Zihao Lin, Hehe Fan
The 3D geometry of real-world scene data is often incomplete. Mainstream methods use depth estimators to inpaint missing structure. However, their prediction results can be inconsi…
Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning
Haoyuan Li, Zhengdong Hu, Jun Wang +2
This paper explores agentic 3D spatial understanding, i.e., MLLM agents performing 3D reasoning through tool use. Existing methods often misuse tools and exhibit biased tool prefer…
One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement
Yixiao Zhou, Dongzhou Cheng, zhiliang wu +3
Large Language Models (LLMs) often fail to utilize their latent reasoning capabilities due to a distributional mismatch between ambiguous human inquiries and the structured logic r…
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
Haoyuan Li, Rui Liu, Hehe Fan +1
Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to learn complex reasoning from long-horizon human interactions. While Multi-modal Large Language Mod…