5 papers · 1 filter
Equilibrium Residuals Expose Three Regimes of Matrix-Game Strategic Reasoning in Language Models
Wenhua Nie, Binhan Luo, Zijie Meng +2
Large language models can score well on named game-theory benchmarks while failing on the same strategic computation once semantic cues are removed. We show this gap with procedura…
Identified-Set Geometry of Distributional Model Extraction under Top- Censored API Access
Wenhua Nie, ZiCheng Zhu, Jianan Wu +3
Modern LLM APIs often reveal only top- logit scores and censor the remaining vocabulary. We study the per-position distribution-recovery limits of this access model. For censori…
Future Validity is the Missing Statistic: From Impossibility to -Estimation for Grammar-Faithful Speculative Decoding
Wenhua Nie, Zijie Meng, Kun Zou +5
Grammar-constrained generation is often combined with local vocabulary masking and speculative decoding, but the resulting sampling law is not the grammar-conditional distribution…
Gradient Starvation in Binary-Reward GRPO: Why Group-Mean Centering Fails and Why the Simplest Fix Works
Wenhua Nie, Jianan Wu, Junlin Liu +6
Group Relative Policy Optimization (GRPO) is a standard algorithm for reinforcement learning from verifiable rewards, but its group-mean-centered advantage can fail under binary re…
The Coupling Tax: How Shared Token Budgets Undermine Visible Chain-of-Thought Under Fixed Output Limits
Wenhua Nie, Junlin Liu, Jianan Wu +5
Chain-of-thought reasoning is often treated as a monotone way to improve language-model accuracy by letting a model think longer. We identify a countervailing effect, the coupling…