9 papers
UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets
Xian Li, Rong Wei, Lujie Yang +6
The paper presents UniPhys, a framework that automatically converts raw 3D models into simulation-ready assets with unified physical semantics, and UniPhysGen, a model that jointly…
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
Zhengxi Lu, Zhiyuan Yao, Jinyang Wu +7
Agent skills, structured packages of procedural knowledge and executable resources that agents dynamically load at inference time, have become a reliable mechanism for augmenting L…
PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
Haoyu Zheng, Yun Zhu, Yuqian Yuan +4
Strategic planning is critical for multi-step reasoning, yet compact Large Language Models (LLMs) often lack the capacity to formulate global strategies, leading to error propagati…
ReMemNav: A Rethinking and Memory-Augmented Framework for Zero-Shot Object Navigation
Feng Wu, Wei Zuo, Wenliang Yang +3
Zero-shot object navigation requires agents to locate unseen target objects in unfamiliar environments without prior maps or task-specific training which remains a significant chal…
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
Junbin Xiao, Shenglang Zhang, Pengxiang Zhu +1
We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the came…
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
Keyu Wang, Bingchen Miao, Wendong Bu +7
The development of Multimodal Virtual Agents has made significant progress through the integration of Multimodal Large Language Models. However, mainstream training paradigms face…