3 papers
cs.CV2026
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
Tianpeng Bu, Xin Liu, Qihua Chen +7
While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge this gap at both the evaluation…
cs.CL2026
Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
Junlin He, Yihong Tang, Tong Nie +5
Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high training efficiency. Although th…
cs.LG2026
Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization
Hao Jiang, Shurui Li, Tianpeng Bu +7
Recent advances in online reinforcement learning (RL) for large language models (LLMs) have demonstrated promising performance in complex reasoning tasks. However, they often exhib…