activity
20242026
most citedUNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG2026

Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation

Dong Nie

Large language model post-training methods such as supervised fine-tuning (SFT), reinforcement learning (RL), and distillation are often analyzed through their loss functions: maxi…

cs.LG2026

Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models

Kai Zhao, Dongliang Nie, Yuchen Lin +4

Joint-Embedding Predictive Architectures (JEPAs) provide a simpleframework for learning world models by predicting future latent representations.However, JEPA training is subject t…

cs.LG20261 cited

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types

Zhichao Wang, Bin Bi, Can Huang +7

RL alignment methods, including RLHF and DPO, are primarily based on pairwise preference data. Although scalar or score-based feedback has been collected in some settings, it is ra…

cs.CV2026

Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models

Siyou Li, Huanan Wu, Juexi Shao +10

Despite the recent advances in the video understanding ability of multimodal large language models (MLLMs), long video understanding remains a challenge. One of the main issues is…

cs.LG2025

Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation

Siyou Li, Pengyao Qin, Huanan Wu +4

Automated radiology report generation (RRG) aims to produce detailed textual reports from clinical imaging, such as computed tomography (CT) scans, to improve the accuracy and effi…

eess.IV2024

ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation

Siyou Li, Beining Xu, Yihao Luo +2

Automatic medical report generation (MRG), which aims to produce detailed text reports from medical images, has emerged as a critical task in this domain. MRG systems can enhance r…