4 papers
Large Language Model Evaluation via Matrix Nuclear-Norm
Yahan Li, Tingyu Xia, Yi Chang +1
As large language models (LLMs) continue to evolve, efficient evaluation metrics are vital for assessing their ability to compress information and reduce redundancy. While traditio…
Length-Controlled Margin-Based Preference Optimization without Reference Model
Gengxu Li, Tingyu Xia, Yi Chang +1
Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF), designed to improve training simp…
A Survey of RWKV
Zhiyuan Li, Tingyu Xia, Yi Chang +1
The Receptance Weighted Key Value (RWKV) model offers a novel alternative to the Transformer architecture, merging the benefits of recurrent and attention-based systems. Unlike con…
Rethinking Data Selection at Scale: Random Selection is Almost All You Need
Tingyu Xia, Bowen Yu, Kai Dang +5
Supervised fine-tuning (SFT) is crucial for aligning Large Language Models (LLMs) with human instructions. The primary goal during SFT is to select a small yet representative subse…