2 papers
cs.LG2026
Learning More from Less: Reinforcement Learning from Hindsight
Iris Xu, Sunshine Jiang, John Marangola +8
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, ma…
cs.LG2026
BOAD: Discovering Hierarchical Software Engineering Agents via Bandit Optimization
Iris Xu, Guangtao Zeng, Zexue He +5
Large language models (LLMs) have shown strong reasoning and coding capabilities, yet they struggle to generalize to real-world software engineering (SWE) problems that are long-ho…