1 citations · 2 across the 4 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CL2026
SteerRM: Debiasing Reward Models via Sparse Autoencoders
Mengyuan Sun, Zhuohao Yu, Weizheng Gu +2
Reward models (RMs) are critical components of alignment pipelines, yet they exhibit biases toward superficial stylistic cues, preferring better-presented responses over semantical…
cs.LG2026
What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?
Weizheng Gu, Chengze Li, Zhuohao Yu +6
Large language models are increasingly evaluated as interactive agents, yet standard agent benchmarks conflate two qualitatively distinct sources of success: semantic tool-use and…