From the 1 of 5 linked papers with an AI index.
5 papers
ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning
Yilun Kong, Yunpeng Qing, Guozheng Ma +4
The paper proposes ExToken, a framework that conditions vision‑language‑action policies on discrete behavioral tokens derived from offline demonstrations to promote diverse, struct…
QPO: Query-dependent Prompt Optimization via Multi-Loop Offline Reinforcement Learning
Yilun Kong, Hangyu Mao, Qi Zhao +7
Prompt engineering has demonstrated remarkable success in enhancing the performance of large language models (LLMs) across diverse tasks. However, most existing prompt optimization…
Decision Flow Policy Optimization
Jifeng Hu, Sili Huang, Siyuan Guo +6
In recent years, generative models have shown remarkable capabilities across diverse fields, including images, videos, language, and decision-making. By applying powerful generativ…
Neuron-level Balance between Stability and Plasticity in Deep Reinforcement Learning
Jiahua Lan, Sen Zhang, Haixia Pan +3
In contrast to the human ability to continuously acquire knowledge, agents struggle with the stability-plasticity dilemma in deep reinforcement learning (DRL), which refers to the…
Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal
Jifeng Hu, Li Shen, Sili Huang +5
Artificial neural networks, especially recent diffusion-based models, have shown remarkable superiority in gaming, control, and QA systems, where the training tasks' datasets are u…