12 papers
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
Mingxuan Zheng, Yujin Zhou, Chuxue Cao +6
LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the a…
Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment
Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang +16
Zero2Skill is a robot learning system that autonomously collects, verifies, and resets manipulation data while using a large language model to store and reuse human corrections, dr…
A Control Theory of Predictability in Latent World Models
Hanzhe You, Yonggang Zhang, Maohao Ran +6
Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward. Current…
CaveAgent: Transforming LLMs into Stateful Runtime Operators
Maohao Ran, Zhenglin Wan, Cooper Lin +21
LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that struggle with long-horizon tasks…
NVV-SuperBench: Beyond Words, Beyond Quality-Benchmarking Nonverbal Vocalizations in Speech Generation
Liumeng Xue, Weizhen Bian, Jiahao Pan +9
Nonverbal vocalizations (NVVs), such as laughing, sighing, and sobbing, are essential for human-like speech, yet standardized evaluation rarely jointly assesses whether systems gen…
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models
Zhe Qian, Yanbiao Ma, Zhuohan Ouyang +7
Multimodal Large Reasoning Models (MLRMs) have achieved remarkable strides in visual reasoning through test time compute scaling, yet long chain reasoning remains prone to hallucin…