2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2026
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
Chengjun Pan, Shichun Liu, Jiahang Lin +10
Reinforcement learning (RL) for LLM post-training faces a fundamental design choice: whether to use a learned critic as a baseline for policy optimization. Classical theory favors…
cs.RO2025★ 2 cited
Online Imitation Learning for Manipulation via Decaying Relative Correction through Teleoperation
Cheng Pan, Hung Hon Cheng, Josie Hughes
Teleoperated robotic manipulators enable the collection of demonstration data, which can be used to train control policies through imitation learning. However, such methods can req…