papers

Publications (45)

cs.CL2026

Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions

Luyang Fang, Xiaowei Yu, Jiazhang Cai +23

The exponential growth of Large Language Models (LLMs) continues to highlight the need for efficient strategies to meet ever-expanding computational and data demands. This survey p…

math.ST2020

Asymptotic Analysis of Sampling Estimators for Randomized Numerical Linear Algebra Algorithms

Ping Ma, Xinlian Zhang, Xin Xing +2

The statistical analysis of Randomized Numerical Linear Algebra (RandNLA) algorithms within the past few years has mostly focused on their performance as point estimators. However,…

cond-mat.mes-hall2019

Waveguide-integrated van der Waals heterostructure photodetector at telecom band with high speed and high responsivity

Nikolaus Flöry, Ping Ma, Yannick Salamin +5

Intensive efforts have been devoted to exploit novel optoelectronic devices based on two-dimensional (2D) transition-metal dichalcogenides (TMDCs) owing to their strong light-matte…

cs.CV2025

DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging

Huimin Cheng, Xiaowei Yu, Shushan Wu +7

Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work l…

stat.ME2020

An Asympirical Smoothing Parameters Selection Approach for Smoothing Spline ANOVA Models in Large Samples

Xiaoxiao Sun, Wenxuan Zhong, Ping Ma

Large samples have been generated routinely from various sources. Classic statistical models, such as smoothing spline ANOVA models, are not well equipped to analyze such large sam…

stat.CO2016

Smoothing spline ANOVA for super-large samples: Scalable computation via rounding parameters

Nathaniel E. Helwig, Ping Ma

In the current era of big data, researchers routinely collect and analyze data of super-large sample sizes. Data-oriented statistical methods have been developed to extract informa…

cs.LG2026

A Single Revision Step Improves Token-Efficient LLM Reasoning

Yingchuan Zhang, Terry Ma, Wenxuan Zhong +1

Large language models (LLMs) achieve higher accuracy on challenging reasoning tasks by scaling test-time compute through multiple trajectory sampling. However, standard aggregation…

stat.ME2013

A Statistical Perspective on Algorithmic Leveraging

Ping Ma, Michael W. Mahoney, Bin Yu

One popular method for dealing with large-scale data sets is sampling. For example, by using the empirical statistical leverage scores as an importance sampling distribution, the m…

stat.ML2023

Optimal Sampling Designs for Multi-dimensional Streaming Time Series with Application to Power Grid Sensor Data

Rui Xie, Shuyang Bai, Ping Ma

The Internet of Things (IoT) system generates massive high-speed temporally correlated streaming data and is often connected with online inference tasks under computational or ener…

stat.ME2026

Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors

Luyang Fang, Yongkai Chen, Jiazhang Cai +2

Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. Ho…

stat.ME2022

Smoothing splines approximation using Hilbert curve basis selection

Cheng Meng, Jun Yu, Yongkai Chen +2

Smoothing splines have been used pervasively in nonparametric regressions. However, the computational burden of smoothing splines is significant when the sample size is large.…

stat.CO2020

More efficient approximation of smoothing splines via space-filling basis selection

Cheng Meng, Xinlian Zhang, Jingyi Zhang +2

We consider the problem of approximating smoothing spline estimators in a nonparametric regression model. When applied to a sample of size , the smoothing spline estimator can b…

physics.optics2018

Plasmonically enhanced graphene photodetector featuring 100 GBd, high-responsivity and compact size

Ping Ma, Yannick Salamin, Benedikt Baeuerle +5

Graphene has shown great potentials for high-speed photodetection. Yet, the responsivities of graphene-based high-speed photodetectors are commonly limited by the weak effective ab…

stat.ME2017

Adaptive Basis Selection for Exponential Family Smoothing Splines with Application in Joint Modeling of Multiple Sequencing Samples

Ping Ma, Nan Zhang, Jianhua Z. Huang +1

Second-generation sequencing technologies have replaced array-based technologies and become the default method for genomics and epigenomics analysis. Second-generation sequencing t…

physics.app-ph2019

Plasmonic Ferroelectric Modulators

Andreas Messner, Felix Eltes, Ping Ma +7

Integrated ferroelectric plasmonic modulators featuring large bandwidths, broad optical operation range, resilience to high temperature and ultracompact footprint are introduced. M…

physics.optics2018

Single Atom Plasmonic Switch

Alexandros Emboras, Jens Niegemann, Ping Ma +5

The atom sets an ultimate scaling limit to Moores law in the electronics industry. And while electronics research already explores atomic scales devices, photonics research still d…

cs.LG2021

Sufficient dimension reduction for classification using principal optimal transport direction

Cheng Meng, Jun Yu, Jingyi Zhang +2

Sufficient dimension reduction is used pervasively as a supervised dimension reduction approach. Most existing sufficient dimension reduction methods are developed for data with a…

stat.ML2021

A Review on Modern Computational Optimal Transport Methods with Applications in Biomedical Research

Jingyi Zhang, Wenxuan Zhong, Ping Ma

Optimal transport has been one of the most exciting subjects in mathematics, starting from the 18th century. As a powerful tool to transport between two probability measures, optim…

stat.ME2021

Best Subset Selection: Statistical Computing Meets Quantum Computing

Wenxuan Zhong, Yuan Ke, Ye Wang +3

With the rapid development of quantum computers, quantum algorithms have been studied extensively. However, quantum algorithms tackling statistical problems are still lacking. In t…

stat.ML2021

Large-scale optimal transport map estimation using projection pursuit

Cheng Meng, Yuan Ke, Jingyi Zhang +3

This paper studies the estimation of large-scale optimal transport maps (OTM), which is a well-known challenging problem owing to the curse of dimensionality. Existing literature a…

cs.AI2026

Confidence-Aware Automated Assessment of Student-Drawn Scientific Models

Luyang Fang, Yingchuan Zhang, Jongchan Park +3

Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standa…

q-bio.QM2025

Large Language Models for Bioinformatics

Wei Ruan, Yanjun Lyu, Jing Zhang +52

With the rapid advancements in large language model (LLM) technology and the emergence of bioinformatics-specific language models (BioLMs), there is a growing need for a comprehens…

stat.ME2021

Minimax Nonparametric Two-sample Test under Smoothing

Xin Xing, Zuofeng Shang, Pang Du +3

We consider the problem of comparing probability densities between two groups. A new probabilistic tensor product smoothing spline framework is developed to model the joint density…

stat.CO2026

Quantum Statistical Bootstrap

Yongkai Chen, Ping Ma, Wenxuan Zhong

The bootstrap is a foundational tool in statistical inference, but its classical implementation relies on Monte Carlo resampling, introducing approximation error and incurring high…

cs.LG2025

Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts

Luyang Fang, Tao Wang, Ping Ma +1

Automated scoring of written constructed responses typically relies on separate models per task, straining computational resources, storage, and maintenance in real-world education…

cs.AI2025

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

Haoran Lu, Luyang Fang, Ruidong Zhang +47

Due to the remarkable capabilities and growing impact of large language models (LLMs), they have been deeply integrated into many aspects of society. Thus, ensuring their alignment…

cs.CV2024

Non-Destructive Peat Analysis using Hyperspectral Imaging and Machine Learning

Yijun Yan, Jinchang Ren, Barry Harrison +3

Peat, a crucial component in whisky production, imparts distinctive and irreplaceable flavours to the final product. However, the extraction of peat disrupts ancient ecosystems and…

eess.IV2025

S2MNet: Speckle-To-Mesh Net for Three-Dimensional Cardiac Morphology Reconstruction via Echocardiogram

Xilin Gong, Yongkai Chen, Shushan Wu +3

Echocardiogram is the most commonly used imaging modality in cardiac assessment duo to its non-invasive nature, real-time capability, and cost-effectiveness. Despite its advantages…

cs.CL2024

Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments

Ehsan Latif, Luyang Fang, Ping Ma +1

This study proposes a method for knowledge distillation (KD) of fine-tuned Large Language Models (LLMs) into smaller, more efficient, and accurate neural networks. We specifically…

stat.ME2025

Online Sequential Leveraging Sampling Method for Streaming Autoregressive Time Series with Application to Seismic Data

Rui Xie, T. N. Sriram, Wei Biao Wu +1

Seismic data contain complex temporal information that arrives at high speed and has a large, even potentially unbounded volume. The explosion of temporally correlated streaming da…

stat.ME2019

Optimal Penalized Function-on-Function Regression under a Reproducing Kernel Hilbert Space Framework

Xiaoxiao Sun, Pang Du, Xiao Wang +1

Many scientific studies collect data where the response and predictor variables are both functions of time, location, or some other covariate. Understanding the relationship betwee…

stat.ME2015

Optimal Subsampling Approaches for Large Sample Linear Regression

Rong Zhu, Ping Ma, Michael W. Mahoney +1

A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample…

stat.ME2026

Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation

Luyang Fang, Haoran Lu, Yongkai Chen +2

As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge distillation reduces deployment cos…

stat.ME2023

Leverage classifier: Another look at support vector machine

Yixin Han, Jun Yu, Nan Zhang +4

Support vector machine (SVM) is a popular classifier known for accuracy, flexibility, and robustness. However, its intensive computation has hindered its application to large-scale…

stat.ME2026

Wahkon: A Statistically Principled Deep RKHS Superposition Network

Yongkai Chen, Wenxuan Zhong, Ping Ma

Deep learning excels at prediction but often lacks finite-sample guarantees and calibrated uncertainty; RKHS (Reproducing Kernel Hilbert Space)-based methods provide those guarante…

cs.CV2025

Semi-supervised Concept Bottleneck Models

Lijie Hu, Tianhao Huang, Huanyi Xie +6

Concept Bottleneck Models (CBMs) have garnered increasing attention due to their ability to provide concept-based explanations for black-box deep learning models while achieving hi…

cs.CL2025

Efficient Multi-Task Inferencing: Model Merging with Gromov-Wasserstein Feature Alignment

Luyang Fang, Ehsan Latif, Haoran Lu +3

Automatic scoring of student responses enhances efficiency in education, but deploying a separate neural network for each task increases storage demands, maintenance efforts, and r…

cs.CY2024

A Systematic Assessment of OpenAI o1-Preview for Higher Order Thinking in Education

Ehsan Latif, Yifan Zhou, Shuchen Guo +24

As artificial intelligence (AI) continues to advance, it demonstrates capabilities comparable to human intelligence, with significant potential to transform education and workforce…

stat.ME2020

LowCon: A design-based subsampling approach in a misspecified linear modeL

Cheng Meng, Rui Xie, Abhyuday Mandal +3

We consider a measurement constrained supervised learning problem, that is, (1) full sample of the predictors are given; (2) the response observations are unavailable and expensive…

cs.AI2026

NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning

Haoran Lu, Luyang Fang, Wenxuan Zhong +1

Multi-agent language systems are often built as hand-designed workflows, where agents are assigned semantic roles and communication protocols are specified in advance. We propose N…

stat.ME2008

Penalized Clustering of Large Scale Functional Data with Multiple Covariates

Ping Ma, Wenxuan Zhong

In this article, we propose a penalized clustering method for large scale data with multiple covariates through a functional data approach. In the proposed method, responses and co…

stat.ML2022

An optimal transport approach for selecting a representative subsample with application in efficient kernel density estimation

Jingyi Zhang, Cheng Meng, Jun Yu +3

Subsampling methods aim to select a subsample as a surrogate for the observed sample. Such methods have been used pervasively in large-scale data analytics, active learning, and pr…

cs.LG2021

Managing dataset shift by adversarial validation for credit scoring

Hongyi Qian, Baohui Wang, Ping Ma +3

Dataset shift is common in credit scoring scenarios, and the inconsistency between the distribution of training data and the data that actually needs to be predicted is likely to c…

stat.CO2018

Optimal Subsampling for Large Sample Logistic Regression

HaiYing Wang, Rong Zhu, Ping Ma

For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary…

math.ST2005

Optimal smoothing in nonparametric mixed-effect models

Chong Gu, Ping Ma

Mixed-effect models are widely used for the analysis of correlated data such as longitudinal data and repeated measures. In this article, we study an approach to the nonparametric…