activity
20242026
collaborators

7 papers

cs.CL2026

MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation

Pham Khanh Chi, Quoc Phong Dao, Thuat Nguyen +3

Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers or token-level outputs, igno…

cs.CL2026

MIPIC: Matryoshka Representation Learning via Self-Distilled Intra-Relational and Progressive Information Chaining

Phung Gia Huy, Hai An Vu, Minh-Phuc Truong +4

Representation learning is fundamental to NLP, but building embeddings that work well at different computational budgets is challenging. Matryoshka Representation Learning (MRL) of…

cs.CL2026

CTPD: Cross Tokenizer Preference Distillation

Truong Nguyen, Phi Van Dat, Ngan Nguyen +3

While knowledge distillation has seen widespread use in pre-training and instruction tuning, its application to aligning language models with human preferences remains underexplore…

cs.LG2025

Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning

The Viet Bui, Tien Mai, Hong Thanh Nguyen

We study the problem of online multi-agent reinforcement learning (MARL) in environments with sparse rewards, where reward feedback is not provided at each interaction but only rev…

cs.LG2025

MisoDICE: Multi-Agent Imitation from Unlabeled Mixed-Quality Demonstrations

The Viet Bui, Tien Mai, Hong Thanh Nguyen

We study offline imitation learning (IL) in cooperative multi-agent settings, where demonstrations have unlabeled mixed quality - containing both expert and suboptimal trajectories…

cs.LG2025

O-MAPL: Offline Multi-agent Preference Learning

The Viet Bui, Tien Mai, Hong Thanh Nguyen

Inferring reward functions from demonstrations is a key challenge in reinforcement learning (RL), particularly in multi-agent RL (MARL), where large joint state-action spaces and c…