4 papers
How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models
Hongxing Wang, Harenome Razanajato, Zhen Zhang +2
Long-context prefill remains expensive because full/GQA layers still score the historical sequence, even in hybrid models with local, sparse, linear, or recurrent components. We st…
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
Shihang Zhang, Mingjin Kuai, Ye Wei +2
Video Moment Retrieval (VMR) task requires accurately localizing temporal boundaries aligned with natural language queries, but many models suffer from a misalignment between conti…
Exploring Training and Inference Scaling Laws in Generative Retrieval
Hongru Cai, Yongqi Li, Ruifeng Yuan +4
Generative retrieval reformulates retrieval as an autoregressive generation task, where large language models (LLMs) generate target documents directly from a query. As a novel par…
Unsupervised Query Routing for Retrieval Augmented Generation
Feiteng Mu, Liwen Zhang, Yong Jiang +4
Query routing for retrieval-augmented generation aims to assign an input query to the most suitable search engine. Existing works rely heavily on supervised datasets that require e…