activity
20242026
collaborators

8 papers

cs.LG2026

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge

Pengxiao Lin, Zheng-An Chen, Zhi-Qin John Xu

Large Language Models (LLMs) excel at multi-hop reasoning in distribution, yet fail on unseen compositions, a phenomenon known as the curse of two-hop reasoning. In this work, we a…

cs.LG2026

Focus and Dilution: The Multi-stage Learning Process of Attention

Zheng-An Chen, Pengxiao Lin, Zhi-Qin John Xu +1

Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identif…

cs.LG2026

Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data

Tianyi Chen, Pengxiao Lin, Zhiwei Wang +1

State Space Models (SSMs) have emerged as promising alternatives to attention mechanisms, with the Mamba architecture demonstrating impressive performance and linear complexity for…

cs.LG2025

Scalable Complexity Control Facilitates Reasoning Ability of LLMs

Liangkai Hang, Junjie Yao, Zhiwei Bai +17

The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their…

cs.CL2025

Reasoning Bias of Next Token Prediction Training

Pengxiao Lin, Zhongwang Zhang, Zhi-Qin John Xu

Since the inception of Large Language Models (LLMs), the quest to efficiently train them for superior reasoning capabilities has been a pivotal challenge. The dominant training par…

cs.CL2025

Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers

Zhongwang Zhang, Pengxiao Lin, Zhiwei Wang +2

Transformers have demonstrated impressive capabilities across various tasks, yet their performance on compositional problems remains a subject of debate. In this study, we investig…