Showing cs.IRShow all
2 papers · 1 filter
cs.IR2026
Do We Need Bigger Models for Science? Task-Aware Retrieval with Small Language Models
Florian Kelber, Matthias Jobst, Yuni Susanti +1
Scientific knowledge discovery increasingly relies on large language models, yet many existing scholarly assistants depend on proprietary systems with tens or hundreds of billions…
cs.IR2026
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
Cornelius Kummer, Lena Jurkschat, Michael Färber +1
With the wide adoption of language models for IR -- and specifically RAG systems -- the latency of the underlying LLM becomes a crucial bottleneck, since the long contexts of retri…