5 papers
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
Ruichen Xu, Wenjing Yan, Ying-Jun Angela Zhang
Understanding reasoning in large language models is complicated by evaluations that conflate multiple reasoning types. We isolate analogical reasoning, where a model transfers an a…
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
Zhiyuan Zhai, Xinkai You, Wenjing Yan +1
Reasoning-capable large language models solve hard problems by emitting long chains of thought, paying heavily in latency, GPU time, and energy. Casual inspection of their traces r…
Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis
Zhiyuan Zhai, Wenjing Yan, Xiaodan Shao +1
Does reinforcement learning genuinely expand what LLM agents can do, or merely make them more reliable? For static reasoning, recent work answers the second: base and RL pass@k cur…
FISMO: Fisher-Structured Momentum-Orthogonalized Optimizer
Chenrui Xu, Wenjing Yan, Ying-Jun Angela Zhang
Training large-scale neural networks requires solving nonconvex optimization where the choice of optimizer fundamentally determines both convergence behavior and computational effi…
Problem-Parameter-Free Decentralized Bilevel Optimization
Zhiwei Zhai, Wenjing Yan, Ying-Jun Angela Zhang
Decentralized bilevel optimization has garnered significant attention due to its critical role in solving large-scale machine learning problems. However, existing methods often rel…