4 citations · 5 across the 5 of their papers we have counts for
4 papers · 1 filter
Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority
Zhanming Shen, Zeyu Qin, Jiaqi Hu +7
The transition from fitting empirical data to achieving true human utility is fundamentally constrained by a granularity mismatch, where fine-grained autoregressive generation is o…
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
Zhanming Shen, Hao Chen, Yulei Tang +6
Instruction tuning is vital for aligning large language models (LLMs) with human intent, but current methods typically rely on costly human-annotated seed data or powerful external…
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
Hao Chen, Haoze Li, Zhiqing Xiao +6
Aligning general-purpose large language models (LLMs) to downstream tasks often incurs significant training adjustment costs. Prior research has explored various avenues to enhance…
LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization
Qi Zhang, Shouqing Yang, Lirong Gao +8
Large language models (LLMs) have demonstrated impressive capabilities in reasoning with the emergence of reasoning models like OpenAI-o1 and DeepSeek-R1. Recent research focuses o…