7 papers
MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation
Pham Khanh Chi, Quoc Phong Dao, Thuat Nguyen +3
Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers or token-level outputs, igno…
MIPIC: Matryoshka Representation Learning via Self-Distilled Intra-Relational and Progressive Information Chaining
Phung Gia Huy, Hai An Vu, Minh-Phuc Truong +4
Representation learning is fundamental to NLP, but building embeddings that work well at different computational budgets is challenging. Matryoshka Representation Learning (MRL) of…
CTPD: Cross Tokenizer Preference Distillation
Truong Nguyen, Phi Van Dat, Ngan Nguyen +3
While knowledge distillation has seen widespread use in pre-training and instruction tuning, its application to aligning language models with human preferences remains underexplore…
Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning
The Viet Bui, Tien Mai, Hong Thanh Nguyen
We study the problem of online multi-agent reinforcement learning (MARL) in environments with sparse rewards, where reward feedback is not provided at each interaction but only rev…
MisoDICE: Multi-Agent Imitation from Unlabeled Mixed-Quality Demonstrations
The Viet Bui, Tien Mai, Hong Thanh Nguyen
We study offline imitation learning (IL) in cooperative multi-agent settings, where demonstrations have unlabeled mixed quality - containing both expert and suboptimal trajectories…
O-MAPL: Offline Multi-agent Preference Learning
The Viet Bui, Tien Mai, Hong Thanh Nguyen
Inferring reward functions from demonstrations is a key challenge in reinforcement learning (RL), particularly in multi-agent RL (MARL), where large joint state-action spaces and c…