123 citations · 306 across the 26 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
Haoran He, Yuxiao Ye, Qingpeng Cai +4
RL with Verifiable Rewards (RLVR) has emerged as a promising paradigm for improving the reasoning abilities of large language models (LLMs). Current methods rely primarily on polic…
cs.LG2025
Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding
StepFun, :, Bin Wang +195
Large language models (LLMs) face low hardware efficiency during decoding, especially for long-context reasoning tasks. This paper introduces Step-3, a 321B-parameter VLM with hard…
cs.LG2022★ 9 cited
Task-Specific Expert Pruning for Sparse Mixture-of-Experts
Tianyu Chen, Shaohan Huang, Yuan Xie +5
The sparse Mixture-of-Experts (MoE) model is powerful for large-scale pre-training and has achieved promising results due to its model capacity. However, with trillions of paramete…