activity
20182025
most citedDeepFlame: A deep learning empowered open-source platform for reacting flow simulations

70 citations · 105 across the 27 of their papers we have counts for

collaborators

35 papers

cs.LG2025

Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models

Junjie Yao, Zhi-Qin John Xu

The embedding space of language models is widely believed to capture the semantic relationships; for instance, embeddings of digits often exhibit an ordered structure that correspo…

cs.LG2025

An overview of condensation phenomenon in deep learning

Zhi-Qin John Xu, Yaoyu Zhang, Zhangchen Zhou

In this paper, we provide an overview of a common phenomenon, condensation, observed during the nonlinear training of neural networks: During the nonlinear training of neural netwo…

cs.CL2025

Reasoning Bias of Next Token Prediction Training

Pengxiao Lin, Zhongwang Zhang, Zhi-Qin John Xu

Since the inception of Large Language Models (LLMs), the quest to efficiently train them for superior reasoning capabilities has been a pivotal challenge. The dominant training par…

cs.CL2025

An Analysis for Reasoning Bias of Language Models with Small Initialization

Junjie Yao, Zhongwang Zhang, Zhi-Qin John Xu

Transformer-based Large Language Models (LLMs) have revolutionized Natural Language Processing by demonstrating exceptional performance across diverse tasks. This study investigate…

math.NA2025

On understanding and overcoming spectral biases of deep neural network learning methods for solving PDEs

Zhi-Qin John Xu, Lulu Zhang, Wei Cai

In this review, we survey the latest approaches and techniques developed to overcome the spectral bias towards low frequency of deep neural network learning methods in learning mul…

cs.LG2024

A rationale from frequency perspective for grokking in training neural network

Zhangchen Zhou, Yaoyu Zhang, Zhi-Qin John Xu

Grokking is the phenomenon where neural networks NNs initially fit the training data and later generalize to the test data during training. In this paper, we empirically provide a…