27 papers
SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance
Yufei Zhang, Chenlu Zhan, Donghui Sun +2
Affordance grounding aims to localize the functional region for interaction, such as the handle to grasp or the button to press, rather than the whole object. This makes it more ch…
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
Yunjie Chen, Xiaoxin Chen, Fang Wang
Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models…
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
Kangsheng Duan, Ziyang Xu, Wenyu Liu +3
While 10B-level industrial foundation models have pushed the boundaries of image inpainting, their prohibitive computational costs severely hinder practical deployment. Constructin…
SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning
Jichao Wang, Liuyang Bian, Yufeng Zhou +9
As Multimodal Large Language Models (MLLMs) mature, GUI agents are evolving from static interactions to complex navigation. While Reinforcement Learning (RL) has emerged as a promi…
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Hao Wang, Guozhi Wang, Han Xiao +8
Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited by sparse rewards and long hori…
Benchmark Shadows: Data Alignment, Parameter Footprints, and Generalization in Large Language Models
Hongjian Zou, Yidan Wang, Qi Ding +2
Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in…