2 papers
cs.RO2026
UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
Yufei Jia, Zhanxiang Cao, Mingrui Yu +48
Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning are placed on a single GPU-ce…
cs.LG2026
On-Policy Replay for Continual Supervised Fine-Tuning
Yan Chen, Taojie Zhu, Meng Zhang +4
Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers from catastrophic forgetting…