3 papers
cs.CL2025
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
Stephen Zhang, Mustafa Khan, Vardan Papyan
Large language models (LLMs) often concentrate their attention on a few specific tokens referred to as attention sinks. Common examples include the first token, a prompt-independen…
cs.CL2024
Enhancing Knowledge Distillation for LLMs with Response-Priming Prompting
Vijay Goyal, Mustafa Khan, Aprameya Tirupati +3
Large language models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing (NLP) tasks. However, these models are often difficult to d…
cs.IR2024
Multi-Aspect Reviewed-Item Retrieval via LLM Query Decomposition and Aspect Fusion
Anton Korikov, George Saad, Ethan Baron +3
While user-generated product reviews often contain large quantities of information, their utility in addressing natural language product queries has been limited, with a key challe…