18 papers
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
Lechao Zhang, Haoran Xu, Jingyu Gong +3
Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limi…
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards
Wentao Yan, Shengqin Wang, Huichi Zhou +4
Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outcome-based supervision and th…
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
Shengqin Wang, Wentao Yan, Huichi Zhou +4
Agentic multimodal models have garnered significant attention for their ability to leverage external tools to tackle complex tasks. However, it is observed that such agents often m…
Memory Intelligence Agent
Jingyang Qiao, Weicheng Meng, Yu Cheng +6
Deep research agents (DRAs) integrate LLM reasoning with external tools. Memory systems enable DRAs to leverage historical experiences, which are essential for efficient reasoning…
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
Xiaofan Li, Ming Yang, Zhiyuan Ma +9
Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs). However, effectively managin…
DailyArt: Discovering Articulation from Single Static Images via Latent Dynamics
Hang Zhang, Qijian Tian, Jingyu Gong +4
Articulated objects are essential for embodied AI and world models, yet inferring their kinematics from a single closed-state image remains challenging because crucial motion cues…