6 papers
PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
Yichuan Wang, Zhifei Li, Zirui Wang +5
Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipe…
ModernVBERT: Towards Smaller Visual Document Retrievers
Paul Teiletche, Quentin Macé, Max Conti +4
Retrieving specific information from a large corpus of documents is a prevalent industrial use case of modern AI, notably due to the popularity of Retrieval-Augmented Generation (R…
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100
We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…
Revisiting Multilingual Data Mixtures in Language Model Pretraining
Negar Foroutan, Paul Teiletche, Ayush Kumar Tarun +1
The impact of different multilingual data mixtures in pretraining large language models (LLMs) has been a topic of ongoing debate, often raising concerns about potential trade-offs…
MMORE: Massive Multimodal Open RAG & Extraction
Alexandre Sallinen, Stefan Krsteski, Paul Teiletche +7
We introduce MMORE, an open-source pipeline for Massive Multimodal Open RetrievalAugmented Generation and Extraction, designed to ingest, transform, and retrieve knowledge from het…
Enhancing Inflation Nowcasting with LLM: Sentiment Analysis on News
Marc-Antoine Allard, Paul Teiletche, Adam Zinebi
This study explores the integration of large language models (LLMs) into classic inflation nowcasting frameworks, particularly in light of high inflation volatility periods such as…