activity
20242026
most citedA Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Simply Stabilizing the Loop via Fully Looped Transformer

Rao Fu, Zixuan Yang, Jiankun Zhang +4

Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same Transformer blocks, trading a…

cs.LG2026

A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning

Siyuan Guo, Yanchao Sun, Jifeng Hu +5

Offline reinforcement learning (RL) provides a promising solution to learning an agent fully relying on a data-driven paradigm. However, constrained by the limited quality of the o…

cs.LG2025

Decision Flow Policy Optimization

Jifeng Hu, Sili Huang, Siyuan Guo +6

In recent years, generative models have shown remarkable capabilities across diverse fields, including images, videos, language, and decision-making. By applying powerful generativ…

cs.LG2025

Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning

Jifeng Hu, Sili Huang, Zhejian Yang +6

Conditional decision generation with diffusion models has shown powerful competitiveness in reinforcement learning (RL). Recent studies reveal the relation between energy-function-…

cs.LG2025

Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal

Jifeng Hu, Li Shen, Sili Huang +5

Artificial neural networks, especially recent diffusion-based models, have shown remarkable superiority in gaming, control, and QA systems, where the training tasks' datasets are u…

cs.LG2024

Solving Continual Offline RL through Selective Weights Activation on Aligned Spaces

Jifeng Hu, Sili Huang, Li Shen +7

Continual offline reinforcement learning (CORL) has shown impressive ability in diffusion-based lifelong learning systems by modeling the joint distributions of trajectories. Howev…