5 papers
Select to Think: Unlocking SLM Potential with Local Sufficiency
Wenxuan Ye, Yangyang Zhang, Xueli An +2
Small language models (SLMs) offer efficient deployment, yet they often lag behind their larger counterparts (LLMs) in reasoning. Existing remedies either invoke an LLM at points o…
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
Jinze Li, Yang Zhang, Xin Yang +5
Autonomous LLM agents increasingly operate in long-horizon, interactive settings where success depends on reusing experience accumulated over extended histories. However, existing…
Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
Yihang Yao, Guangtao Zeng, Raina Wu +4
Reinforcement learning (RL) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs). While RL has demonstrated substantial perfo…
CommVQ: Commutative Vector Quantization for KV Cache Compression
Junyan Li, Yang Zhang, Muhammad Yusuf Hassan +8
Large Language Models (LLMs) are increasingly used in applications requiring long context lengths, but the key-value (KV) cache often becomes a memory bottleneck on GPUs as context…
Steering LLM Thinking with Budget Guidance
Junyan Li, Wenshuo Zhao, Yang Zhang +1
Recent deep-thinking large language models often reason extensively to improve performance, but such lengthy reasoning is not always desirable, as it incurs excessive inference cos…