papers

Publications (18)

cs.LG2025

Rethinking Benign Overfitting in Two-Layer Neural Networks

Ruichen Xu, Kexin Chen

Recent theoretical studies (Kou et al., 2023; Cao et al., 2022) have revealed a sharp phase transition from benign to harmful overfitting when the noise-to-feature ratio exceeds a…

cs.LG2026

Provable Knowledge Acquisition and Extraction in One-Layer Transformers

Ruichen Xu, Kexin Chen

Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. Despite growing empirical evidence that MLP lay…

cs.CL2026

Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers

Ruichen Xu, Wenjing Yan, Ying-Jun Angela Zhang

Understanding reasoning in large language models is complicated by evaluations that conflate multiple reasoning types. We isolate analogical reasoning, where a model transfers an a…

cs.LG2026

Differential Privacy in Two-Layer Networks: How DP-SGD Harms Fairness and Robustness

Ruichen Xu, Kexin Chen

Differentially private learning is essential for training models on sensitive data, but empirical studies consistently show that it can degrade performance, introduce fairness issu…

math.NT2025

The Gan-Gross-Prasad period of Klingen Eisenstein families over unitary groups

Ruichen Xu

In this article, we compute the Gan-Gross-Prasad period integral of Klingen Eisenstein series over the unitary group with a cuspidal automorphic form over $\…

cs.LG2025

Kolmogorov-Arnold Representation for Symplectic Learning: Advancing Hamiltonian Neural Networks

Zongyu Wu, Ruichen Xu, Luoyao Chen +3

We propose a Kolmogorov-Arnold Representation-based Hamiltonian Neural Network (KAR-HNN) that replaces the Multilayer Perceptrons (MLPs) with univariate transformations. While Hami…

cs.LG2026

Towards Principled Test-Time Adaptation for Time Series Forecasting

Haochun Wang, Ruichen Xu, Georgios Kementzidis +3

Test-time adaptation (TTA) has recently emerged as a promising approach for improving time series forecasting (TSF) under distribution shift. Existing TSF-TTA methods differ in how…

math.NT2025

On the Mordell-Weil rank of certain CM abelian varieties over anticyclotomic towers

Haidong Li, Ruichen Xu

Let be an imaginary quadratic extension, and let be an odd prime. In this paper, we investigate the growth of Mordell-Weil ranks of CM abelian varieties associat…

physics.comp-ph2025

Velocity-Inferred Hamiltonian Neural Networks: Learning Energy-Conserving Dynamics from Position-Only Data

Ruichen Xu, Zongyu Wu, Luoyao Chen +5

Data-driven modeling of physical systems often relies on learning both positions and momenta to accurately capture Hamiltonian dynamics. However, in many practical scenarios, only…

physics.comp-ph2025

Boundary-Informed Method of Lines for Physics Informed Neural Networks

Maximilian Cederholm, Siyao Wang, Haochun Wang +2

We propose a hybrid solver that fuses the dimensionality-reduction strengths of the Method of Lines (MOL) with the flexibility of Physics-Informed Neural Networks (PINNs). Instead…

cs.RO2024

Safe Hybrid-Action Reinforcement Learning-Based Decision and Control for Discretionary Lane Change

Ruichen Xu, Xiao Liu, Jinming Xu +1

Autonomous lane-change, a key feature of advanced driver-assistance systems, can enhance traffic efficiency and reduce the incidence of accidents. However, safe driving of autonomo…

cs.LG2026

Tackling Privacy Heterogeneity in Differentially Private Federated Learning

Ruichen Xu, Ying-Jun Angela Zhang, Jianwei Huang

Differentially private federated learning (DP-FL) enables clients to collaboratively train machine learning models while preserving the privacy of their local data. However, most e…

cs.LG2026

JSAM: Privacy Straggler-Resilient Joint Client Selection and Incentive Mechanism Design in Differentially Private Federated Learning

Ruichen Xu, Ying-Jun Angela Zhang, Jianwei Huang

Differentially private federated learning faces a fundamental tension: privacy protection mechanisms that safeguard client data simultaneously create quantifiable privacy costs tha…

cs.LG2026

GAGA: Gaussianity-Aware Gaussian Approximation for Efficient 3D Molecular Generation

Jingxiang Qu, Wenhan Gao, Ruichen Xu +1

Gaussian Probability Path based Generative Models (GPPGMs) generate data by reversing a stochastic process that progressively corrupts samples with Gaussian noise. Despite state-of…

cs.LG2026

HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning

Ruichen Xu, Jingxiang Qu, Wenhan Gao +5

Graph self-supervised learning aims to learn transferable representations from large-scale unlabeled graph data. Joint-embedding predictive architectures (JEPAs) avoid explicit neg…

math.OC2025

The Impact of Move Schemes on Simulated Annealing Performance

Ruichen Xu, Haochun Wang, Yuefan Deng

Designing an effective move-generation function for Simulated Annealing (SA) in complex models remains a significant challenge. In this work, we present a combination of theoretica…

cs.LG2025

Can Data-Driven Dynamics Reveal Hidden Physics? There Is A Need for Interpretable Neural Operators

Wenhan Gao, Jian Luo, Fang Wan +4

Recently, neural operators have emerged as powerful tools for learning mappings between function spaces, enabling data-driven simulations of complex dynamics. Despite their success…

cs.LG2025

An Iterative Framework for Generative Backmapping of Coarse Grained Proteins

Georgios Kementzidis, Erin Wong, John Nicholson +2

The techniques of data-driven backmapping from coarse-grained (CG) to fine-grained (FG) representation often struggle with accuracy, unstable training, and physical realism, especi…