3 papers
cs.CL2026
Can LLMs Take Retrieved Information with a Grain of Salt?
Behzad Shayegh, Mohamed Osama Ahmed, Fred Tung +1
Large language models have demonstrated impressive retrieval-augmented capabilities. However, a crucial area remains underexplored: their ability to appropriately adapt responses t…
cs.LG2025
Radar: Fast Long-Context Decoding for Any Transformer
Yongchang Hao, Mengyao Zhai, Hossein Hajimirsadeghi +2
Transformer models have demonstrated exceptional performance across a wide range of applications. Though forming the foundation of Transformer models, the dot-product attention doe…
cs.LG2024
Were RNNs All We Needed?
Leo Feng, Frederick Tung, Mohamed Osama Ahmed +2
The introduction of Transformers in 2017 reshaped the landscape of deep learning. Originally proposed for sequence modelling, Transformers have since achieved widespread success ac…