Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
Jie Hao, Rui Yu, Wei Zhang +3
Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream tasks. However, existing approac…
cs.LG2025
Reward Models in Deep Reinforcement Learning: A Survey
Rui Yu, Shenghua Wan, Yucen Wang +4
In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are intr…