4 papers
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
Yanshi Li, Xueru Bai, Shuman Liu +1
The paper introduces RepBench, a framework that aggregates thousands of benchmark datasets into a large set of probe texts to evaluate capability-aligned representations of large l…
Decomposing and Steering Functional Metacognition in Large Language Models
Yanshi Li, Xueru Bai, Shuman Liu +2
Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies in benchmark settings. Prior…
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
Zhenpeng Su, Leiyu Pan, Xue Bai +8
We present Klear-Reasoner, a model with long reasoning capabilities that demonstrates careful deliberation during problem solving, achieving outstanding performance across multiple…
MiLe Loss: a New Entropy-Weighed Loss for Mitigating the Bias of Learning Difficulties in Large Language Models
Zhenpeng Su, Xing Wu, Xue Bai +5
Generative language models are usually pretrained on large text corpus via predicting the next token (i.e., sub-word/word/phrase) given the previous ones. Recent works have demonst…