70 citations · 105 across the 27 of their papers we have counts for
35 papers
Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models
Junjie Yao, Zhi-Qin John Xu
The embedding space of language models is widely believed to capture the semantic relationships; for instance, embeddings of digits often exhibit an ordered structure that correspo…
An overview of condensation phenomenon in deep learning
Zhi-Qin John Xu, Yaoyu Zhang, Zhangchen Zhou
In this paper, we provide an overview of a common phenomenon, condensation, observed during the nonlinear training of neural networks: During the nonlinear training of neural netwo…
Reasoning Bias of Next Token Prediction Training
Pengxiao Lin, Zhongwang Zhang, Zhi-Qin John Xu
Since the inception of Large Language Models (LLMs), the quest to efficiently train them for superior reasoning capabilities has been a pivotal challenge. The dominant training par…
An Analysis for Reasoning Bias of Language Models with Small Initialization
Junjie Yao, Zhongwang Zhang, Zhi-Qin John Xu
Transformer-based Large Language Models (LLMs) have revolutionized Natural Language Processing by demonstrating exceptional performance across diverse tasks. This study investigate…
On understanding and overcoming spectral biases of deep neural network learning methods for solving PDEs
Zhi-Qin John Xu, Lulu Zhang, Wei Cai
In this review, we survey the latest approaches and techniques developed to overcome the spectral bias towards low frequency of deep neural network learning methods in learning mul…
A rationale from frequency perspective for grokking in training neural network
Zhangchen Zhou, Yaoyu Zhang, Zhi-Qin John Xu
Grokking is the phenomenon where neural networks NNs initially fit the training data and later generalize to the test data during training. In this paper, we empirically provide a…