4 papers
Parallel Context-of-Experts Decoding for Retrieval Augmented Generation
Giulio Corallo, Paolo Papotti
Retrieval Augmented Generation faces a trade-off: concatenating documents in a long prompt enables multi-document reasoning but creates prefill bottlenecks, while encoding document…
A Word is Worth 4-bit: Efficient Log Parsing with Binary Coded Decimal Recognition
Prerak Srivastava, Giulio Corallo, Sergey Rybalko
System-generated logs are typically converted into categorical log templates through parsing. These templates are crucial for generating actionable insights in various downstream t…
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
Giulio Corallo, Orion Weller, Fabio Petroni +1
Incorporating external knowledge in large language models (LLMs) enhances their utility across diverse applications, but existing methods have trade-offs. Retrieval-Augmented Gener…
Latent Abstractions in Generative Diffusion Models
Giulio Franzese, Mattia Martini, Giulio Corallo +2
In this work we study how diffusion-based generative models produce high-dimensional data, such as an image, by implicitly relying on a manifestation of a low-dimensional set of la…