2 papers
cs.LG2026
SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning
Jichao Wang, Liuyang Bian, Yufeng Zhou +9
As Multimodal Large Language Models (MLLMs) mature, GUI agents are evolving from static interactions to complex navigation. While Reinforcement Learning (RL) has emerged as a promi…
cs.LG2026
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Hao Wang, Guozhi Wang, Han Xiao +8
Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited by sparse rewards and long hori…