Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures
Xiyin Zeng, Yuyu Sun, Haoyang Li +2
Vision-Language-Action systems follow instructions to execute multi-step tasks in multimodal environments. Recent VLA approaches typically rely on post-hoc correction mechanisms or…
cs.AI2026
RLHFless: Serverless Computing for Efficient RLHF
Rui Wei, Hanfei Yu, Shubham Jain +5
Reinforcement Learning from Human Feedback (RLHF) has been widely applied to Large Language Model (LLM) post-training to align model outputs with human preferences. Recent models,…