Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
Yuxin Chen, Liang Luo, Buyun Zhang +44
Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this wor…
cs.LG2026
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents
Peng Xu, Sijia Chen, Junzhuo Li +1
Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ti…