From the 1 of 4 linked papers with an AI index.
4 papers
Not All Retrievals are Useful: Cross-Attention for Input-Aware RAG in Time Series Forecasting
Seunghan Lee, Jaehoon Lee, Jun Seo +7
The paper introduces Cross-RAG, a retrieval-augmented generation framework for zero-shot time series forecasting that uses query‑retrieval cross‑attention to selectively attend to…
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
Jaehoon Lee, Mingi Jung, Soohyuk Jang +3
Large Vision-Language Models (VLMs) achieve strong multimodal understanding capabilities by leveraging high-resolution visual inputs, but the resulting large number of visual token…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
Training LLMs over Neurally Compressed Text
Brian Lester, Jaehoon Lee, Alex Alemi +4
In this paper, we explore the idea of training large language models (LLMs) over highly compressed text. While standard subword tokenizers compress text by a small factor, neural t…