14 papers
Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools
Xiuwei Chen, Quanlin Chen, Wentao Hu +8
The paper introduces Beyond the Eye (BEE), an implicit visual‑tool framework for multimodal large language models that learns to self‑regulate when to invoke visual tools, reducing…
PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
Bin Hu, Yanwen Ma, Jiehui Huang +14
Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering traje…
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
Xiuwei Chen, Wentao Hu, Hanhui Li +9
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…
Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
Kun Xiang, Terry Jingchen Zhang, Yinya Huang +13
The rapid advancement of embodied intelligence and world models has intensified efforts to integrate physical laws into AI systems, yet physical perception and symbolic physics rea…
ProPhy: Progressive Physical Alignment for Dynamic World Simulation
Zijun Wang, Panwen Hu, Jing Wang +7
Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent resul…
ERGO: Excess-Risk-Guided Optimization for High-Fidelity Monocular 3D Gaussian Splatting
Zehua Ma, Hanhui Li, Zhenyu Xie +4
Generating 3D content from a single image remains a fundamentally challenging and ill-posed problem due to the inherent absence of geometric and textural information in occluded re…