Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs
Xi Chen, Mingyu Jin, Jingcheng Niu +7
In this paper, we present empirical and theoretical evidence against a central but largely implicit assumption in circuit and sheaf discovery (CSD), which we term the Functional An…
cs.CL2025
-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
Yining Wang, Jinman Zhao, Chuangxin Zhao +3
Reinforcement Learning with Human Feedback (RLHF) has been the dominant approach for improving the reasoning capabilities of Large Language Models (LLMs). Recently, Reinforcement L…