activity
20242026
most citedAgentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Simply Stabilizing the Loop via Fully Looped Transformer

Rao Fu, Zixuan Yang, Jiankun Zhang +4

Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same Transformer blocks, trading a…

cs.LG2026

Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning

Rufeng Chen, Zhaofan Zhang, Zhejiang Yang +2

Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a single episode. While diffusi…

cs.LG2025

Decision Flow Policy Optimization

Jifeng Hu, Sili Huang, Siyuan Guo +6

In recent years, generative models have shown remarkable capabilities across diverse fields, including images, videos, language, and decision-making. By applying powerful generativ…

cs.LG2025

Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning

Jifeng Hu, Sili Huang, Zhejian Yang +6

Conditional decision generation with diffusion models has shown powerful competitiveness in reinforcement learning (RL). Recent studies reveal the relation between energy-function-…

cs.LG2024

Solving Continual Offline RL through Selective Weights Activation on Aligned Spaces

Jifeng Hu, Sili Huang, Li Shen +7

Continual offline reinforcement learning (CORL) has shown impressive ability in diffusion-based lifelong learning systems by modeling the joint distributions of trajectories. Howev…

cs.LG2024

Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal

Jifeng Hu, Li Shen, Sili Huang +5

Artificial neural networks, especially recent diffusion-based models, have shown remarkable superiority in gaming, control, and QA systems, where the training tasks' datasets are u…