7 papers
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
Truong Nguyen, Tien-Phat Nguyen, Linh Ngo Van +3
Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences over full sequences even thou…
LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models
Minh Chu Xuan, Tien-Phat Nguyen, Linh Ngo Van +3
Cross-lingual topic modeling aims to discover shared semantic structures across languages, yet existing models depend on sparse bilingual resources and often yield incoherent or we…
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
Tien-Phat Nguyen, Truong Nguyen, Minh-Phuc Truong +3
Muon orthogonalizes the momentum buffer before each update, replacing its singular values with ones via Newton-Schulz iterations. This simple change lets Muon tolerate far larger l…
Selective Off-Policy Reference Tuning with Plan Guidance
Duc Anh Le, Tien-Phat Nguyen, Thien Huu Nguyen +2
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those fa…
BSO: Safety Alignment Is Density Ratio Matching
Tien-Phat Nguyen, Truong Nguyen, Thin Nguyen +3
Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement learning, and primal-dual update…
S-Chain: Structured Visual Chain-of-Thought For Medicine
Khai Le-Duc, Duy M. H. Nguyen, Phuong T. H. Trinh +21
Faithful reasoning in medical vision-language models (VLMs) requires not only accurate predictions but also transparent alignment between textual rationales and visual evidence. Wh…