14 citations · 21 across the 16 of their papers we have counts for
5 papers · 1 filter
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
Mohamed Sana, Nicola Piovesan, Antonio De Domenico +2
We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contain more responses with negative…
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
Dmitry Akulov, Mohamed Sana, Antonio De Domenico +3
Large language models (LLMs) rely on key-value (KV) caches for efficient autoregressive decoding; however, cache size grows linearly with context length and model depth, becoming a…
Goal-Oriented Time-Series Forecasting: Foundation Framework Design
Luca-Andrei Fechete, Mohamed Sana, Fadhel Ayed +4
Conventional time-series forecasting methods typically aim to minimize overall prediction error, without accounting for the varying importance of different forecast ranges in downs…
FlexTrain: A Dynamic Training Framework for Heterogeneous Devices Environments
Mert Unsal, Ali Maatouk, Antonio De Domenico +2
As deep learning models become increasingly large, they pose significant challenges in heterogeneous devices environments. The size of deep learning models makes it difficult to de…
Anomaly Detection at Scale: The Case for Deep Distributional Time Series Models
Fadhel Ayed, Lorenzo Stella, Tim Januschowski +1
This paper introduces a new methodology for detecting anomalies in time series data, with a primary application to monitoring the health of (micro-) services and cloud resources. T…