collaborators

10 papers

cs.AI2026

Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization

Pengfei Xu, Yong Liu, Xiaoya Nan +2

Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction history into closed-loop decisio…

cs.AI2026

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought

Zeyu Gan, Hao Yi, Yong Liu

Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a pivotal paradigm for enhancing the rea…

cs.AI2026

Structure-Conditioned Actor-Critic Branches for Quality-Diversity Reinforcement Learning

Lianrong Zuo, Peilan Xu, Yong Liu +1

Quality-diversity reinforcement learning (QD-RL) aims to construct policy repertoires that contain both high-performing and behaviorally diverse policies. Existing QD-RL methods ma…

cs.CL2026

AsyncLane: Decoupling Refinement from Advancement in Diffusion Language Model Decoding

Yingxuan Ren, Yuxuan Lou, Yong Liu +4

Block-wise semi-autoregressive decoding is the standard inference paradigm for diffusion large language models (DLMs), but it imposes a strict dependency between blocks: the next b…

cs.AI2026

Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents

Zeyu Gan, Huayi Tang, Yong Liu

As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With t…

cs.AI2026

Capabilities and Fundamental Limits of Latent Chain-of-Thought

Jiaxuan Zou, Yaozhong Xiong, Yong Liu

Latent Chain-of-Thought (Latent CoT) models promise efficient reasoning via continuous representations, yet exhibit puzzling performance inconsistencies: excelling at exploration (…