2 papers
cs.LG2025
Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
Huikang Su, Dengyun Peng, Zifeng Zhuang +4
Offline safe reinforcement learning aims to learn policies that satisfy predefined safety constraints from static datasets. Existing sequence-model-based methods condition action g…
cs.CL2025
Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models
Qiguang Chen, Dengyun Peng, Jinhao Liu +4
Recent advancements in large language models (LLMs) have greatly improved their capabilities on complex reasoning tasks through Long Chain-of-Thought (CoT). However, this approach…