3 papers
cs.LG2026
How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models
Hongxing Wang, Harenome Razanajato, Zhen Zhang +2
Long-context prefill remains expensive because full/GQA layers still score the historical sequence, even in hybrid models with local, sparse, linear, or recurrent components. We st…
cs.IR2025
Exploring Training and Inference Scaling Laws in Generative Retrieval
Hongru Cai, Yongqi Li, Ruifeng Yuan +4
Generative retrieval reformulates retrieval as an autoregressive generation task, where large language models (LLMs) generate target documents directly from a query. As a novel par…
cs.IR2025
Unsupervised Query Routing for Retrieval Augmented Generation
Feiteng Mu, Liwen Zhang, Yong Jiang +4
Query routing for retrieval-augmented generation aims to assign an input query to the most suitable search engine. Existing works rely heavily on supervised datasets that require e…