most citedFalcon2-11B Technical Report

2 citations · 3 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2024

Investigating Regularization of Self-Play Language Models

Reda Alami, Abdalgader Abubaker, Mastane Achab +2

This paper explores the effects of various forms of regularization in the context of language model alignment via self-play. While both reinforcement learning from human feedback (…

cs.LG2023

A Risk-Averse Framework for Non-Stationary Stochastic Multi-Armed Bandits

Reda Alami, Mohammed Mahfoud, Mastane Achab

In a typical stochastic multi-armed bandit problem, the objective is often to maximize the expected sum of rewards over some time horizon . While the choice of a strategy that a…

cs.LG2023

Deep Reinforcement Learning Algorithms for Hybrid V2X Communication: A Benchmarking Study

Fouzi Boukhalfa, Reda Alami, Mastane Achab +2

In today's era, autonomous vehicles demand a safety level on par with aircraft. Taking a cue from the aerospace industry, which relies on redundancy to achieve high reliability, th…

cs.LG20231 cited

Emergent Communication in Multi-Agent Reinforcement Learning for Future Wireless Networks

Marwa Chafii, Salmane Naoumi, Reda Alami +3

In different wireless network scenarios, multiple network entities need to cooperate in order to achieve a common task with minimum delay and energy consumption. Future wireless ne…

cs.LG2023

One-Step Distributional Reinforcement Learning

Mastane Achab, Reda Alami, Yasser Abdelaziz Dahou Djilali +2

Reinforcement learning (RL) allows an agent interacting sequentially with an environment to maximize its long-term expected return. In the distributional RL (DistrRL) paradigm, the…

cs.LG2023

Restarted Bayesian Online Change-point Detection for Non-Stationary Markov Decision Processes

Reda Alami, Mohammed Mahfoud, Eric Moulines

We consider the problem of learning in a non-stationary reinforcement learning (RL) environment, where the setting can be fully described by a piecewise stationary discrete-time Ma…