collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

EntroRouter: Learning Efficient Model Routing via Entropy Regulation

Kaiyi Zhang, Xueliang Zhao, Zhuocheng Gong +2

Model routing balances solution accuracy and computational cost by selecting among models of varying capabilities. While recent multi-round frameworks interleave reasoning and plan…

cs.CL2026

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

Jingwen Chen, Wenkai Yang, Shengda Fan +7

Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward continual learning in large l…

cs.CL2026

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning

Yubo Wang, Juntian Zhang, Yichen Wu +3

While Chain-of-Thought empowers Large Vision-Language Models with multi-step reasoning, explicit textual rationales suffer from an information bandwidth bottleneck, where continuou…

cs.CL2025

LaSeR: Reinforcement Learning with Last-Token Self-Rewarding

Wenkai Yang, Weijie Liu, Ruobing Xie +4

Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a core paradigm for enhancing the reasoning capabilities of Large Language Models (LLMs). To address t…

cs.CL2025

Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning

Wenkai Yang, Shuming Ma, Yankai Lin +1

Recent studies have shown that making a model spend more time thinking through longer Chain of Thoughts (CoTs) enables it to gain significant improvements in complex reasoning task…

cs.CL2025

Beyond the Surface: Measuring Self-Preference in LLM Judgments

Zhi-Yuan Chen, Hao Wang, Xinyu Zhang +2

Recent studies show that large language models (LLMs) exhibit self-preference bias when serving as judges, meaning they tend to favor their own responses over those generated by ot…