4 papers · 1 filter
When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
Yuqing Yang, Robin Jia
We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in their previously generated false a…
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Zhen Huang, Zengzhi Wang, Shijie Xia +25
The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showc…
Alignment for Honesty
Yuqing Yang, Ethan Chern, Xipeng Qiu +2
Recent research has made significant strides in aligning large language models (LLMs) with helpfulness and harmlessness. In this paper, we argue for the importance of alignment for…
Weak-to-Strong Reasoning
Yuqing Yang, Yan Ma, Pengfei Liu
When large language models (LLMs) exceed human-level capabilities, it becomes increasingly challenging to provide full-scale and accurate supervision for these models. Weak-to-stro…