6 citations · 10 across the 31 of their papers we have counts for
Showing 2026 · cs.CLShow all
2 papers · 2 filters
cs.CL2026
Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
Jiayun Wu, Peixu Hou, Shan Qu +3
Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior ac…
cs.CL2026
IntPro: A Proxy Agent for Context-Aware Intent Understanding via Retrieval-conditioned Inference
Guanming Liu, Meng Wu, Peng Zhang +8
Large language models (LLMs) have become integral to modern Human-AI collaboration workflows, where accurately understanding user intent serves as a crucial step for generating sat…