Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
VeriBound: PAC-Bayesian Generalization Bounds for Process Reward Models Trained with Formal Verification Tools
Amirul Rahman, Mohammed Sabih Alsharari
Process Reward Models (PRMs) provide step-level verification for Large Language Model (LLM) reasoning, yet their training data acquisition remains a bottleneck: human annotation is…
cs.CL2026
Avoiding Overthinking and Underthinking: Curriculum-Aware Budget Scheduling for LLMs
Amirul Rahman, Aisha Karim, Kenji Nakamura +1
Scaling test-time compute via extended reasoning has become a key paradigm for improving the capabilities of large language models (LLMs). However, existing approaches optimize rea…