3 papers
cs.IR2026
Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design
Bojian Hou, Xiaolong Liu, Xiaoyi Liu +26
Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for designing and allocating resources in massive-s…
cs.CL2025
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
Jiahui Li, Lin Li, Tai-wei Chang +4
Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained…
cs.LG2025
Learning Causal Transition Matrix for Instance-dependent Label Noise
Jiahui Li, Tai-Wei Chang, Kun Kuang +3
Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of lea…