Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
Yaning Jia, Chunhui Zhang, Wenxuan Xu +3
Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. Th…
cs.AI2025
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
Chiyu Ma, Enpei Zhang, Yilun Zhao +7
LLM-as-Judge has emerged as a scalable alternative to human evaluation, enabling large language models (LLMs) to provide reward signals in trainings. While recent work has explored…