3 papers
cs.AI2026
DOPD: Dual On-policy Distillation
Xinlei Yu, Gen Li, Qingyi Si +13
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sourc…
cs.CV2026
RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
Sicheng Feng, Kaiwen Tuo, Song Wang +3
Fine-grained visual reasoning remains a core challenge for multimodal large language models (MLLMs). The recently introduced ReasonMap highlights this gap by showing that even adva…
cs.LG2025
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
Kaiwen Tuo, Huan Wang
State-space language models such as Mamba match Transformer quality while permitting linear complexity inference, yet still comprise billions of parameters that hinder deployment.…