10 papers
Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization
Pengfei Xu, Yong Liu, Xiaoya Nan +2
Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction history into closed-loop decisio…
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought
Zeyu Gan, Hao Yi, Yong Liu
Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a pivotal paradigm for enhancing the rea…
Structure-Conditioned Actor-Critic Branches for Quality-Diversity Reinforcement Learning
Lianrong Zuo, Peilan Xu, Yong Liu +1
Quality-diversity reinforcement learning (QD-RL) aims to construct policy repertoires that contain both high-performing and behaviorally diverse policies. Existing QD-RL methods ma…
AsyncLane: Decoupling Refinement from Advancement in Diffusion Language Model Decoding
Yingxuan Ren, Yuxuan Lou, Yong Liu +4
Block-wise semi-autoregressive decoding is the standard inference paradigm for diffusion large language models (DLMs), but it imposes a strict dependency between blocks: the next b…
Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents
Zeyu Gan, Huayi Tang, Yong Liu
As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With t…
Capabilities and Fundamental Limits of Latent Chain-of-Thought
Jiaxuan Zou, Yaozhong Xiong, Yong Liu
Latent Chain-of-Thought (Latent CoT) models promise efficient reasoning via continuous representations, yet exhibit puzzling performance inconsistencies: excelling at exploration (…