3 papers
math.OC2026
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning
Jialun Cao, Fernando Acero, David Šiška +1
Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigoro…
cs.AI2026
When Do We Need LLMs? A Diagnostic for Language-Driven Bandits
Uljad Berdica, Fernando Acero, Anton Ipsen +3
We study Contextual Multi-Armed Bandits (CMABs) for non-episodic decision-making problems where the context includes both textual and numerical information (e.g., recommendation sy…
cs.AI2026
Distill and Align Decomposition for Enhanced Claim Verification
Jabez Magomere, Elena Kochkina, Samuel Mensah +6
Complex claim verification requires decomposing sentences into verifiable subclaims, yet existing methods struggle to align decomposition quality with verification performance. We…