activity
20242026
collaborators

5 papers

cs.CL2026

CTPD: Cross Tokenizer Preference Distillation

Truong Nguyen, Phi Van Dat, Ngan Nguyen +3

While knowledge distillation has seen widespread use in pre-training and instruction tuning, its application to aligning language models with human preferences remains underexplore…

cs.LG2025

Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning

The Viet Bui, Tien Mai, Hong Thanh Nguyen

We study the problem of online multi-agent reinforcement learning (MARL) in environments with sparse rewards, where reward feedback is not provided at each interaction but only rev…

cs.LG2025

MisoDICE: Multi-Agent Imitation from Unlabeled Mixed-Quality Demonstrations

The Viet Bui, Tien Mai, Hong Thanh Nguyen

We study offline imitation learning (IL) in cooperative multi-agent settings, where demonstrations have unlabeled mixed quality - containing both expert and suboptimal trajectories…

cs.LG2025

O-MAPL: Offline Multi-agent Preference Learning

The Viet Bui, Tien Mai, Hong Thanh Nguyen

Inferring reward functions from demonstrations is a key challenge in reinforcement learning (RL), particularly in multi-agent RL (MARL), where large joint state-action spaces and c…

cs.LG2024

ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift Regularization

The Viet Bui, Thanh Hong Nguyen, Tien Mai

Offline reinforcement learning (RL) has garnered significant attention for its ability to learn effective policies from pre-collected datasets without the need for further environm…