most citedCURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.RO2025

A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning

Shaopeng Zhai, Qi Zhang, Tianyi Zhang +7

Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLA…

cs.LG20251 cited

CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention

Qingbin Li, Rongkun Xue, Jie Wang +8

Recent advances in Reinforcement Learning with Verified Reward (RLVR) have driven the emergence of more sophisticated cognitive behaviors in large language models (LLMs), thereby e…

cs.LG2025

Efficient Skill Discovery via Regret-Aware Optimization

He Zhang, Ming Zhou, Shaopeng Zhai +2

Unsupervised skill discovery aims to learn diverse and distinguishable behaviors in open-ended reinforcement learning. For existing methods, they focus on improving diversity throu…

cs.CL2025

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System

Yuan Guo, Tingjia Miao, Zheng Wu +3

Autonomous agents powered by multimodal large language models have been developed to facilitate task execution on mobile devices. However, prior work has predominantly focused on a…

cs.AI2024

CLSP: High-Fidelity Contrastive Language-State Pre-training for Agent State Representation

Fuxian Huang, Qi Zhang, Shaopeng Zhai +6

With the rapid development of artificial intelligence, multimodal learning has become an important research area. For intelligent agents, the state is a crucial modality to convey…