2 papers
cs.CL2026
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM
Shuvendu Roy, Mengyao Zhai, Hossein Hajimirsadeghi +1
Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-context inputs. However, deploying LLMs is cos…
cs.LG2025
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
Shuvendu Roy, Hossein Hajimirsadeghi, Mengyao Zhai +1
Recent advances in large language models have demonstrated the promise of unsupervised reinforcement learning (RL) methods for enhancing reasoning capabilities without external sup…