10 papers
PReM: Learning What to Preserve and When to Refresh for Context Compression
Bohan Yu, Lei Shen, Chenxi Zhou +5
The paper proposes PReM, a framework that lets language models dynamically decide which parts of a long context to keep and when to refresh stored information, improving efficiency…
OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models
Yuanhao Yue, Chengyu Wang, Yuanjie Lyu +2
Recent multimodal large language models (MLLMs) have shown strong chain-of-thought (CoT) reasoning ability on vision-language tasks, but their direct deployment in real-world syste…
Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards
Yuanjie Lyu, Chengyu Wang, Lei Shen +2
Small LLMs often struggle to match the agentic capabilities of large, costly models. While reinforcement learning can help, progress has been limited by two structural bottlenecks:…
Rewards as Labels: Revisiting RLVR from a Classification Perspective
Zepeng Zhai, Meilin Chen, Jiaxuan Zhao +3
Reinforcement Learning with Verifiable Rewards has recently advanced the capabilities of Large Language Models in complex reasoning tasks by providing explicit rule-based supervisi…
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
Zhuang Yu, Lei Shen, Jing Zhao +1
Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, an…
BARD: budget-aware reasoning distillation
Lujie Niu, Lei Shen, Yi Jiang +4
While long Chain-of-Thought (CoT) distillation effectively transfers reasoning capability to smaller language models, the reasoning process often remains redundant and computationa…