From the 1 of 6 linked papers with an AI index.
6 papers
Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression
Angelo Nardone, Paolo Ferragina
We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and stru…
Extended Depth-First Representations of -trees
Gabriel Carmona, Paolo Ferragina, Giovanni Manzini +1
The paper proposes depth‑first memory layouts for k²‑trees, including plain and balanced‑parentheses variants and compressed versions, and shows they improve cache performance, com…
LLM-based Source Code Compression via Thresholded Symbol Ranking
Angelo Nardone, Paolo Ferragina
We study the problem of lossless compression of source code, motivated by the storage demands of large-scale software archives, such as Software Heritage (https://www.softwareherit…
Recall Before Rerank: Benchmarking Deep Learning Models for Large-Scale Code-to-Code Retrieval
Leonardo Venuta, Francesco Tosoni, Paolo Ferragina
Semantic code search and clone detection are essential for software development, maintenance, and reuse. This paper evaluates the effectiveness, efficiency, and scalability of cont…
Compressibility Measures and Succinct Data Structures for Piecewise Linear Approximations
Paolo Ferragina, Filippo Lari
We study the problem of deriving compressibility measures for Piecewise Linear Approximations (PLAs), i.e., error-bounded approximations of a set of two-dimensional increasing data…
Learned Compression of Nonlinear Time Series With Random Access
Andrea Guerra, Giorgio Vinciguerra, Antonio Boffa +1
Time series play a crucial role in many fields, including finance, healthcare, industry, and environmental monitoring. The storage and retrieval of time series can be challenging d…