activity
20202026
most citedOff-policy Imitation Learning from Visual Inputs

1 citations · 2 across the 5 of their papers we have counts for

collaborators

6 papers

cs.AI2026

A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization

Shiye Lei, Zhihao Cheng, Dacheng Tao

Reinforcement learning (RL) post-training has increasingly demonstrated strong ability to elicit reasoning behaviors in large language models (LLMs). For training efficiency, rollo…

cs.LG2025

Offline Behavioral Data Selection

Shiye Lei, Zhihao Cheng, Dacheng Tao

Behavioral cloning is a widely adopted approach for offline policy learning from expert demonstrations. However, the large scale of offline behavioral datasets often results in com…

cs.LG2025

State Diversity Matters in Offline Behavior Distillation

Shiye Lei, Zhihao Cheng, Dacheng Tao

Offline Behavior Distillation (OBD), which condenses massive offline RL data into a compact synthetic behavioral dataset, offers a promising approach for efficient policy training…

cs.AI20251 cited

Revisiting LLM Reasoning via Information Bottleneck

Shiye Lei, Zhihao Cheng, Kai Jia +1

Large language models (LLMs) have recently demonstrated remarkable progress in reasoning capabilities through reinforcement learning with verifiable rewards (RLVR). By leveraging s…

cs.LG20211 cited

Off-policy Imitation Learning from Visual Inputs

Zhihao Cheng, Li Shen, Dacheng Tao

Recently, various successful applications utilizing expert states in imitation learning (IL) have been witnessed. However, another IL setting -- IL from visual inputs (ILfVI), whic…

cs.RO2020

On the Guaranteed Almost Equivalence between Imitation Learning from Observation and Demonstration

Zhihao Cheng, Liu Liu, Aishan Liu +3

Imitation learning from observation (LfO) is more preferable than imitation learning from demonstration (LfD) due to the nonnecessity of expert actions when reconstructing the expe…