1 paper
Nadav Schneider, Itamar Zimerman, Eliya Nachmani
Sequence models like Transformers and RNNs often overallocate attention to irrelevant context, leading to noisy intermediate representations. This degrades LLM capabilities by prom…