5 papers
AlignSAE: Concept-Aligned Sparse Autoencoders
Minglai Yang, Xinyu Guo, Zhengliang Shi +4
Large Language Models (LLMs) encode factual knowledge within hidden parametric spaces that are difficult to inspect or control. While Sparse Autoencoders (SAEs) can decompose hidde…
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
Minglai Yang, Ethan Huang, Liang Zhang +3
We introduce Grade School Math with Distracting Context (GSM-DC), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically contro…
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
Haris Riaz, Ellen Riloff, Mihai Surdeanu
We propose a simple, unsupervised method that injects pragmatic principles in retrieval-augmented generation (RAG) frameworks such as Dense Passage Retrieval to enhance the utility…
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
Razvan-Gabriel Dumitru, Minglai Yang, Vikas Yadav +1
We introduce CopySpec, a simple yet effective technique to tackle the inefficiencies LLMs face when generating responses that closely resemble previous outputs or responses that ca…
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
Razvan-Gabriel Dumitru, Paul-Ioan Clotan, Vikas Yadav +2
This paper introduces a novel model compression approach through dynamic layer-specific pruning in Large Language Models (LLMs), enhancing the traditional methodology established b…