3 papers
cs.LG2026
Think Wider: Mitigating Latent Rank Collapse in Implicit Chain-of-Thought Reasoning
Yuwen Hao, Menglin Yang
Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing intermediate computation, but explicit rationales increase decoding length,…
cs.SE2026
SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents
Ruitao Wang, Yuwen Hao, Menglin Yang
Web agents often struggle to generalize to unseen websites because they lack website-specific supervision. Recent exploration-based data synthesis methods reduce manual annotation,…
cs.CL2026
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
Jinchang Zhu, Jindong Li, Yuwen Hao +3
A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraining: upper layers commit to s…