6 citations · 7 across the 6 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
Li Wang, Xiaodong Lu, Xiaohan Wang +5
Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…
cs.LG2025
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
Yuyang Xu, Yi Cheng, Haochao Ying +5
Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforceme…
cs.LG2024★ 1 cited
Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
Yi Cheng, Renjun Hu, Haochao Ying +3
Until recently, the question of the effective inductive bias of deep models on tabular data has remained unanswered. This paper investigates the hypothesis that arithmetic feature…