4 papers
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM
Shuvendu Roy, Mengyao Zhai, Hossein Hajimirsadeghi +1
Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-context inputs. However, deploying LLMs is cos…
An Attention-based Model for Robust Forecasting with Missing Modality
Zhitian Zhang, Wenjie Zi, Yunduz Rakhmangulova +3
Learning with missing modalities is a fundamental challenge in multimodal robot learning, as real-world robotic systems often operate in environments with incomplete sensor data. A…
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
Shuvendu Roy, Hossein Hajimirsadeghi, Mengyao Zhai +1
Recent advances in large language models have demonstrated the promise of unsupervised reinforcement learning (RL) methods for enhancing reasoning capabilities without external sup…
Radar: Fast Long-Context Decoding for Any Transformer
Yongchang Hao, Mengyao Zhai, Hossein Hajimirsadeghi +2
Transformer models have demonstrated exceptional performance across a wide range of applications. Though forming the foundation of Transformer models, the dot-product attention doe…