collaborators

6 papers

cs.AI2026

NEX: Neuron Explore-Exploit Scoring for Label-Free Chain-of-Thought Selection and Model Ranking

Kang Chen, Zhuoka Feng, Sihan Zhao +5

Large language models increasingly spend inference compute sampling multiple chain-of-thought traces or searching over merged checkpoints. This shifts the bottleneck from generatio…

cs.AI2026

ARM: Role-Conditioned Neuron Transplantation for Training-Free Generalist LLM Agent Merging

Zhuoka Feng, Kang Chen, Sihan Zhao +7

Interactive large language model agents have advanced rapidly, but most remain specialized to a single environment and fail to adapt robustly to other environments. Model merging o…

cs.CL2025

Do LLMs Signal When They're Right? Evidence from Neuron Agreement

Kang Chen, Yaoning Wang, Kai Xiong +4

Large language models (LLMs) commonly boost reasoning via sample-evaluate-ensemble decoders, achieving label free gains without ground truth. However, prevailing strategies score c…

cs.CL2025

EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization

Yaoning Wang, Jiahao Ying, Yixin Cao +2

The rapid advancement of large language models (LLMs) and the development of increasingly large and diverse evaluation benchmarks have introduced substantial computational challeng…

cs.LG2025

Pruning the Unsurprising: Efficient LLM Reasoning via First-Token Surprisal

Wenhao Zeng, Yaoning Wang, Chao Hu +4

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities by scaling up the length of Chain-of-Thought (CoT). However, excessively long reasoning traces pose substant…

cs.CL2025

Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric

Yixin Cao, Jiahao Ying, Yaoning Wang +3

Large Language Models (LLMs) have become indispensable across academia, industry, and daily applications, yet current evaluation methods struggle to keep pace with their rapid deve…