From the 1 of 5 linked papers with an AI index.
5 papers
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
Zhanzhi Lou, Hui Chen, Yibo Li +2
The paper introduces Meta-TTL, a bi‑level optimization framework that learns adaptation policies for test‑time learning in language agents, using evolutionary search to improve per…
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Qian Wang, Zhanzhi Lou, Zhenheng Tang +2
LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittl…
Making Bias Non-Predictive: Training Robust LLM Reasoning via Reinforcement Learning
Qian Wang, Xuandong Zhao, Zirui Zhang +4
Large language models (LLMs) increasingly serve as reasoners and automated evaluators, yet they remain susceptible to cognitive biases -- often altering their reasoning when faced…
Towards Evaluting Fake Reasoning Bias in Language Models
Qian Wang, Zhenheng Tang, Zhanzhi Lou +3
Large Reasoning Models (LRMs), evolved from standard Large Language Models (LLMs), are increasingly utilized as automated judges because of their explicit reasoning processes. Yet…
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
Qian Wang, Zhanzhi Lou, Zhenheng Tang +5
Large Reasoning Models (LRMs) like DeepSeek-R1 and OpenAI-o1 have demonstrated remarkable reasoning capabilities, raising important questions about their biases in LLM-as-a-judge s…