8 papers
verdi: retrieval is not transfer for continual world model optimization
Junyu Wu, Shiqin Nie, Youyi Kou +9
Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model toward a user-specified objec…
Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling
Sen Cui, Jingheng Ma
World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world…
MetaForge: A Self-Evolving Multimodal Agent that Retrieves, Adapts, and Forges Tools On Demand
Shouang Wei, Houcheng Min, Xinpeng Dong +8
Multimodal agents have achieved notable progress on complex reasoning tasks through tool use, yet remain limited by two issues: statically predefined tool inventories fail to gener…
Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support
Shuying Chen, Sen Cui, Zhong Cao
In this work, we propose Oph-Guid-RAG, a multimodal visual RAG system for ophthalmology clinical question answering and decision support. We treat each guideline page as an indepen…
FATE: Closed-Loop Feasibility-Aware Task Generation with Active Repair for Physically Grounded Robotic Curricula
Bingchuan Wei, Bingqi Huang, Jingheng Ma +2
Recent breakthroughs in generative simulation have harnessed Large Language Models (LLMs) to generate diverse robotic task curricula, yet these open-loop paradigms frequently produ…
Reversible Diffusion Decoding for Diffusion Language Models
Xinyun Wang, Min Zhang, Sen Cui +4
Diffusion language models enable parallel token generation through block-wise decoding, but their irreversible commitments can lead to stagnation, where the reverse diffusion proce…