5 papers
Before Parc Fermé: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving
Luca Benfenati, Ali Azimi, Matteo Risso +3
Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but their memory and generation la…
Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
Luca Benfenati, Matteo Risso, Andrea Vannozzi +5
Key-value (KV) caching enables fast autoregressive decoding but at long contexts becomes a dominant bottleneck in High Bandwidth Memory (HBM) capacity and bandwidth. A common mitig…
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
Lorenz K. Müller, Philippe Bich, Jiawei Zhuang +3
Post-training quantization has emerged as the most widely used strategy for deploying large language models at low precision. Still, current methods show perplexity degradation at…
Foundation Models for Structural Health Monitoring
Luca Benfenati, Daniele Jahier Pagliari, Luca Zanatta +6
Structural Health Monitoring (SHM) is a critical task for ensuring the safety and reliability of civil infrastructures, typically realized on bridges and viaducts by means of vibra…
EnhancePPG: Improving PPG-based Heart Rate Estimation with Self-Supervision and Augmentation
Luca Benfenati, Sofia Belloni, Alessio Burrello +6
Heart rate (HR) estimation from photoplethysmography (PPG) signals is a key feature of modern wearable devices for health and wellness monitoring. While deep learning models show p…