works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.RO2026

ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning

Yilun Kong, Yunpeng Qing, Guozheng Ma +4

The paper proposes ExToken, a framework that conditions vision‑language‑action policies on discrete behavioral tokens derived from offline demonstrations to promote diverse, struct…

cs.AI2025

QPO: Query-dependent Prompt Optimization via Multi-Loop Offline Reinforcement Learning

Yilun Kong, Hangyu Mao, Qi Zhao +7

Prompt engineering has demonstrated remarkable success in enhancing the performance of large language models (LLMs) across diverse tasks. However, most existing prompt optimization…

cs.LG2025

Decision Flow Policy Optimization

Jifeng Hu, Sili Huang, Siyuan Guo +6

In recent years, generative models have shown remarkable capabilities across diverse fields, including images, videos, language, and decision-making. By applying powerful generativ…

cs.AI2025

Neuron-level Balance between Stability and Plasticity in Deep Reinforcement Learning

Jiahua Lan, Sen Zhang, Haixia Pan +3

In contrast to the human ability to continuously acquire knowledge, agents struggle with the stability-plasticity dilemma in deep reinforcement learning (DRL), which refers to the…

cs.LG2025

Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal

Jifeng Hu, Li Shen, Sili Huang +5

Artificial neural networks, especially recent diffusion-based models, have shown remarkable superiority in gaming, control, and QA systems, where the training tasks' datasets are u…