11 papers
Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation
Selen Erkan, Bastian Boll, Kristian Kersting +2
Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific formatting requirements. This esp…
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
Max Henning Höth, Kristian Kersting, Björn Deiseroth +1
Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both contributes to and faithfully…
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
Thomas F Burns, Letitia Parcalabescu, Stephan Wäldchen +5
Scaling data quantity is essential for large language models (LLMs), yet recent findings show that data quality can significantly boost performance and training efficiency. We intr…
Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models
Björn Deiseroth, Björn Deiseroth, Max Henning Höth +3
Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic rather than verifiable evidence -- le…
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
Sebastian Sztwiertnia, Felix Friedrich, Kristian Kersting +2
Pre-training decoder-only language models relies on vast amounts of high-quality data, yet the availability of such data is increasingly reaching its limits. While metadata is comm…
Measuring and Guiding Monosemanticity
Ruben Härle, Felix Friedrich, Manuel Brack +4
There is growing interest in leveraging mechanistic interpretability and controllability to better understand and influence the internal dynamics of large language models (LLMs). H…