3 papers
cs.CL2025
-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
Yining Wang, Jinman Zhao, Chuangxin Zhao +3
Reinforcement Learning with Human Feedback (RLHF) has been the dominant approach for improving the reasoning capabilities of Large Language Models (LLMs). Recently, Reinforcement L…
cs.CL2025
UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter-Efficient Fine-Tuning of Large Models
Xueyan Zhang, Jinman Zhao, Zhifei Yang +4
This paper introduces Uniform Orthogonal Reinitialization Adaptation (UORA), a novel parameter-efficient fine-tuning (PEFT) approach for Large Language Models (LLMs). UORA achieves…
cs.CL2025
Role-Play Paradox in Large Language Models: Reasoning Performance Gains and Ethical Dilemmas
Jinman Zhao, Zifan Qian, Linbo Cao +5
Role-play in large language models (LLMs) enhances their ability to generate contextually relevant and high-quality responses by simulating diverse cognitive perspectives. However,…