works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.CL2026

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

Angelo Nardone, Paolo Ferragina

We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and stru…

cs.DS2026

Extended Depth-First Representations of -trees

Gabriel Carmona, Paolo Ferragina, Giovanni Manzini +1

The paper proposes depth‑first memory layouts for k²‑trees, including plain and balanced‑parentheses variants and compressed versions, and shows they improve cache performance, com…

cs.IT2026

LLM-based Source Code Compression via Thresholded Symbol Ranking

Angelo Nardone, Paolo Ferragina

We study the problem of lossless compression of source code, motivated by the storage demands of large-scale software archives, such as Software Heritage (https://www.softwareherit…

cs.SE2026

Recall Before Rerank: Benchmarking Deep Learning Models for Large-Scale Code-to-Code Retrieval

Leonardo Venuta, Francesco Tosoni, Paolo Ferragina

Semantic code search and clone detection are essential for software development, maintenance, and reuse. This paper evaluates the effectiveness, efficiency, and scalability of cont…

cs.DS2025

Compressibility Measures and Succinct Data Structures for Piecewise Linear Approximations

Paolo Ferragina, Filippo Lari

We study the problem of deriving compressibility measures for Piecewise Linear Approximations (PLAs), i.e., error-bounded approximations of a set of two-dimensional increasing data…

cs.LG2024

Learned Compression of Nonlinear Time Series With Random Access

Andrea Guerra, Giorgio Vinciguerra, Antonio Boffa +1

Time series play a crucial role in many fields, including finance, healthcare, industry, and environmental monitoring. The storage and retrieval of time series can be challenging d…