9 papers
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
Yihan Wang, Zhong Guan, Haoran Sun +3
Small language models are attractive backbones for interactive agents, but direct distillation from strong teacher trajectories often turns rich multi-turn behavior into one-shot i…
Robust Contrastive Graph Clustering with Adaptive Local-Global Integration
Lei Zhang, Fubo Sun, Haipeng Yang +2
Graph clustering is essential in graph analysis for revealing structural patterns and node communities. Despite recent advances in self-supervised contrastive learning that have im…
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
Zhong Guan, Yongjian Guo, Haoran Sun +5
Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a c…
RL-VLA: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
Haoran Sun, Yongjian Guo, Zhong Guan +13
Reinforcement learning (RL) has emerged as a critical paradigm for post-training Vision-Language-Action (VLA) models, enabling embodied agents to adapt and improve through environm…
Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration
Zhong Guan, Likang Wu, Hongke Zhao +2
Many existing studies have achieved significant improvements in the reasoning capabilities of large language models (LLMs) through reinforcement learning with verifiable rewards (R…
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
Zhong Guan, Likang Wu, Hongke Zhao +2
Attention mechanisms are critical to the success of large language models (LLMs), driving significant advancements in multiple fields. However, for graph-structured data, which req…