Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Congchao Wang, Diwakar Singh, Qiaozi Gao +3
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement lear…
cs.AI2026
Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
Yaopei Zeng, Congchao Wang, JianHang Chen +3
LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger…
cs.AI2026
ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing
Yaopei Zeng, Congchao Wang, Blake JianHang Chen +1
Routing has emerged as a promising strategy for balancing performance and cost in large language model (LLM) systems that combine lightweight models with powerful but expensive lar…