2 papers
cs.CL2026
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
Hongze Tan, Zihan Wang, Jianfei Pan +7
Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained cr…
cs.AI2025
OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
Debdeep Sanyal, Umakanta Maharana, Yash Sinha +4
Role-based access control (RBAC) and hierarchical structures are foundational to how information flows and decisions are made within virtually all organizations. As the potential o…