3 papers
cs.AI2026
CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning
Ayoub Belouadah, Sylvain Kubler, Yves Le Traon
Safe reinforcement learning (Safe RL) aims to maximize expected return while satisfying safety constraints, typically modeled as Constrained Markov Decision Processes (CMDPs). Whil…
cs.LG2026
Optimized Federated Knowledge Distillation with Distributed Neural Architecture Search
Chaimaa Medjadji, Sylvain Kubler, Yves Le Traon +3
Federated Learning (FL) enables collaborative model training without centralizing data. However, real-world deployments must simultaneously address statistical heterogeneity across…
cs.CL2026
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation
Yu Pan, Yang Hou, Xiongfei Wu +4
Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performance. However, existing S2ST sys…