Publications (18)
Rethinking Benign Overfitting in Two-Layer Neural Networks
Ruichen Xu, Kexin Chen
Recent theoretical studies (Kou et al., 2023; Cao et al., 2022) have revealed a sharp phase transition from benign to harmful overfitting when the noise-to-feature ratio exceeds a…
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
Ruichen Xu, Kexin Chen
Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. Despite growing empirical evidence that MLP lay…
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
Ruichen Xu, Wenjing Yan, Ying-Jun Angela Zhang
Understanding reasoning in large language models is complicated by evaluations that conflate multiple reasoning types. We isolate analogical reasoning, where a model transfers an a…
Differential Privacy in Two-Layer Networks: How DP-SGD Harms Fairness and Robustness
Ruichen Xu, Kexin Chen
Differentially private learning is essential for training models on sensitive data, but empirical studies consistently show that it can degrade performance, introduce fairness issu…
The Gan-Gross-Prasad period of Klingen Eisenstein families over unitary groups
Ruichen Xu
In this article, we compute the Gan-Gross-Prasad period integral of Klingen Eisenstein series over the unitary group with a cuspidal automorphic form over $\…
Kolmogorov-Arnold Representation for Symplectic Learning: Advancing Hamiltonian Neural Networks
Zongyu Wu, Ruichen Xu, Luoyao Chen +3
We propose a Kolmogorov-Arnold Representation-based Hamiltonian Neural Network (KAR-HNN) that replaces the Multilayer Perceptrons (MLPs) with univariate transformations. While Hami…
Towards Principled Test-Time Adaptation for Time Series Forecasting
Haochun Wang, Ruichen Xu, Georgios Kementzidis +3
Test-time adaptation (TTA) has recently emerged as a promising approach for improving time series forecasting (TSF) under distribution shift. Existing TSF-TTA methods differ in how…
On the Mordell-Weil rank of certain CM abelian varieties over anticyclotomic towers
Haidong Li, Ruichen Xu
Let be an imaginary quadratic extension, and let be an odd prime. In this paper, we investigate the growth of Mordell-Weil ranks of CM abelian varieties associat…
Velocity-Inferred Hamiltonian Neural Networks: Learning Energy-Conserving Dynamics from Position-Only Data
Ruichen Xu, Zongyu Wu, Luoyao Chen +5
Data-driven modeling of physical systems often relies on learning both positions and momenta to accurately capture Hamiltonian dynamics. However, in many practical scenarios, only…
Boundary-Informed Method of Lines for Physics Informed Neural Networks
Maximilian Cederholm, Siyao Wang, Haochun Wang +2
We propose a hybrid solver that fuses the dimensionality-reduction strengths of the Method of Lines (MOL) with the flexibility of Physics-Informed Neural Networks (PINNs). Instead…
Safe Hybrid-Action Reinforcement Learning-Based Decision and Control for Discretionary Lane Change
Ruichen Xu, Xiao Liu, Jinming Xu +1
Autonomous lane-change, a key feature of advanced driver-assistance systems, can enhance traffic efficiency and reduce the incidence of accidents. However, safe driving of autonomo…
Tackling Privacy Heterogeneity in Differentially Private Federated Learning
Ruichen Xu, Ying-Jun Angela Zhang, Jianwei Huang
Differentially private federated learning (DP-FL) enables clients to collaboratively train machine learning models while preserving the privacy of their local data. However, most e…
JSAM: Privacy Straggler-Resilient Joint Client Selection and Incentive Mechanism Design in Differentially Private Federated Learning
Ruichen Xu, Ying-Jun Angela Zhang, Jianwei Huang
Differentially private federated learning faces a fundamental tension: privacy protection mechanisms that safeguard client data simultaneously create quantifiable privacy costs tha…
GAGA: Gaussianity-Aware Gaussian Approximation for Efficient 3D Molecular Generation
Jingxiang Qu, Wenhan Gao, Ruichen Xu +1
Gaussian Probability Path based Generative Models (GPPGMs) generate data by reversing a stochastic process that progressively corrupts samples with Gaussian noise. Despite state-of…
HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning
Ruichen Xu, Jingxiang Qu, Wenhan Gao +5
Graph self-supervised learning aims to learn transferable representations from large-scale unlabeled graph data. Joint-embedding predictive architectures (JEPAs) avoid explicit neg…
The Impact of Move Schemes on Simulated Annealing Performance
Ruichen Xu, Haochun Wang, Yuefan Deng
Designing an effective move-generation function for Simulated Annealing (SA) in complex models remains a significant challenge. In this work, we present a combination of theoretica…
Can Data-Driven Dynamics Reveal Hidden Physics? There Is A Need for Interpretable Neural Operators
Wenhan Gao, Jian Luo, Fang Wan +4
Recently, neural operators have emerged as powerful tools for learning mappings between function spaces, enabling data-driven simulations of complex dynamics. Despite their success…
An Iterative Framework for Generative Backmapping of Coarse Grained Proteins
Georgios Kementzidis, Erin Wong, John Nicholson +2
The techniques of data-driven backmapping from coarse-grained (CG) to fine-grained (FG) representation often struggle with accuracy, unstable training, and physical realism, especi…