calibration 1experience replay 1kl regularization 1LLM agents 1LoRA adapters 1model routing 1multimodal llms 1probe routing 1self-evolving critic 1step-level confidence 1training-free methods 1
From the 2 of 5 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Congchao Wang, Diwakar Singh, Qiaozi Gao +3
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement lear…
cs.AI2026
Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
Yaopei Zeng, Congchao Wang, JianHang Chen +3
The paper proposes the Critic Experience Bank, a training-free framework that lets large language model agents estimate confidence for each action by storing and retrieving past st…
cs.AI2026
ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing
Yaopei Zeng, Congchao Wang, Blake JianHang Chen +1
The paper proposes two methods—a attention‑based probe and a KL‑regularized LoRA probe (ReLope)—to improve routing decisions in multimodal large language models by extracting more…