16 papers
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Xiaoying Zhang, Yipeng Zhang, Hao Sun +4
Recent advances in reinforcement learning (RL) using numerical rewards have significantly enhanced the complex reasoning capabilities of large language models (LLMs). However, we i…
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
Jiachen Ma, Jiawen Zhang, Xiangtian Li +3
While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-…
VRPRM: Process Reward Modeling via Visual Reasoning
Xinquan Chen, Chongying Yue, Bangwei Liu +3
Process Reward Model (PRM) is widely used in the post-training of Large Language Model (LLM) because it can perform fine-grained evaluation of the reasoning steps of generated cont…
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
Sirui Chen, Lei Xu, Yuying Zhao +6
Recent RL methods have substantially improved the reasoning abilities of LLMs. Existing reward designs mainly follow two paradigms: (1) Reinforcement learning with verifiable rewar…
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
Yan Xie, Changkui Mao, Changsong Wu +34
As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility…
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
Yuanfu Wang, Zhixuan Liu, Xiangtian Li +2
The prevailing paradigm for training large reasoning models--combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR)--is fundamentally con…