Publications (45)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
Luyang Fang, Xiaowei Yu, Jiazhang Cai +23
The exponential growth of Large Language Models (LLMs) continues to highlight the need for efficient strategies to meet ever-expanding computational and data demands. This survey p…
Asymptotic Analysis of Sampling Estimators for Randomized Numerical Linear Algebra Algorithms
Ping Ma, Xinlian Zhang, Xin Xing +2
The statistical analysis of Randomized Numerical Linear Algebra (RandNLA) algorithms within the past few years has mostly focused on their performance as point estimators. However,…
Waveguide-integrated van der Waals heterostructure photodetector at telecom band with high speed and high responsivity
Nikolaus Flöry, Ping Ma, Yannick Salamin +5
Intensive efforts have been devoted to exploit novel optoelectronic devices based on two-dimensional (2D) transition-metal dichalcogenides (TMDCs) owing to their strong light-matte…
DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging
Huimin Cheng, Xiaowei Yu, Shushan Wu +7
Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work l…
An Asympirical Smoothing Parameters Selection Approach for Smoothing Spline ANOVA Models in Large Samples
Xiaoxiao Sun, Wenxuan Zhong, Ping Ma
Large samples have been generated routinely from various sources. Classic statistical models, such as smoothing spline ANOVA models, are not well equipped to analyze such large sam…
Smoothing spline ANOVA for super-large samples: Scalable computation via rounding parameters
Nathaniel E. Helwig, Ping Ma
In the current era of big data, researchers routinely collect and analyze data of super-large sample sizes. Data-oriented statistical methods have been developed to extract informa…
A Single Revision Step Improves Token-Efficient LLM Reasoning
Yingchuan Zhang, Terry Ma, Wenxuan Zhong +1
Large language models (LLMs) achieve higher accuracy on challenging reasoning tasks by scaling test-time compute through multiple trajectory sampling. However, standard aggregation…
A Statistical Perspective on Algorithmic Leveraging
Ping Ma, Michael W. Mahoney, Bin Yu
One popular method for dealing with large-scale data sets is sampling. For example, by using the empirical statistical leverage scores as an importance sampling distribution, the m…
Optimal Sampling Designs for Multi-dimensional Streaming Time Series with Application to Power Grid Sensor Data
Rui Xie, Shuyang Bai, Ping Ma
The Internet of Things (IoT) system generates massive high-speed temporally correlated streaming data and is often connected with online inference tasks under computational or ener…
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
Luyang Fang, Yongkai Chen, Jiazhang Cai +2
Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. Ho…
Smoothing splines approximation using Hilbert curve basis selection
Cheng Meng, Jun Yu, Yongkai Chen +2
Smoothing splines have been used pervasively in nonparametric regressions. However, the computational burden of smoothing splines is significant when the sample size is large.…
More efficient approximation of smoothing splines via space-filling basis selection
Cheng Meng, Xinlian Zhang, Jingyi Zhang +2
We consider the problem of approximating smoothing spline estimators in a nonparametric regression model. When applied to a sample of size , the smoothing spline estimator can b…
Plasmonically enhanced graphene photodetector featuring 100 GBd, high-responsivity and compact size
Ping Ma, Yannick Salamin, Benedikt Baeuerle +5
Graphene has shown great potentials for high-speed photodetection. Yet, the responsivities of graphene-based high-speed photodetectors are commonly limited by the weak effective ab…
Adaptive Basis Selection for Exponential Family Smoothing Splines with Application in Joint Modeling of Multiple Sequencing Samples
Ping Ma, Nan Zhang, Jianhua Z. Huang +1
Second-generation sequencing technologies have replaced array-based technologies and become the default method for genomics and epigenomics analysis. Second-generation sequencing t…
Plasmonic Ferroelectric Modulators
Andreas Messner, Felix Eltes, Ping Ma +7
Integrated ferroelectric plasmonic modulators featuring large bandwidths, broad optical operation range, resilience to high temperature and ultracompact footprint are introduced. M…
Single Atom Plasmonic Switch
Alexandros Emboras, Jens Niegemann, Ping Ma +5
The atom sets an ultimate scaling limit to Moores law in the electronics industry. And while electronics research already explores atomic scales devices, photonics research still d…
Sufficient dimension reduction for classification using principal optimal transport direction
Cheng Meng, Jun Yu, Jingyi Zhang +2
Sufficient dimension reduction is used pervasively as a supervised dimension reduction approach. Most existing sufficient dimension reduction methods are developed for data with a…
A Review on Modern Computational Optimal Transport Methods with Applications in Biomedical Research
Jingyi Zhang, Wenxuan Zhong, Ping Ma
Optimal transport has been one of the most exciting subjects in mathematics, starting from the 18th century. As a powerful tool to transport between two probability measures, optim…
Best Subset Selection: Statistical Computing Meets Quantum Computing
Wenxuan Zhong, Yuan Ke, Ye Wang +3
With the rapid development of quantum computers, quantum algorithms have been studied extensively. However, quantum algorithms tackling statistical problems are still lacking. In t…
Large-scale optimal transport map estimation using projection pursuit
Cheng Meng, Yuan Ke, Jingyi Zhang +3
This paper studies the estimation of large-scale optimal transport maps (OTM), which is a well-known challenging problem owing to the curse of dimensionality. Existing literature a…
Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
Luyang Fang, Yingchuan Zhang, Jongchan Park +3
Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standa…
Large Language Models for Bioinformatics
Wei Ruan, Yanjun Lyu, Jing Zhang +52
With the rapid advancements in large language model (LLM) technology and the emergence of bioinformatics-specific language models (BioLMs), there is a growing need for a comprehens…
Minimax Nonparametric Two-sample Test under Smoothing
Xin Xing, Zuofeng Shang, Pang Du +3
We consider the problem of comparing probability densities between two groups. A new probabilistic tensor product smoothing spline framework is developed to model the joint density…
Quantum Statistical Bootstrap
Yongkai Chen, Ping Ma, Wenxuan Zhong
The bootstrap is a foundational tool in statistical inference, but its classical implementation relies on Monte Carlo resampling, introducing approximation error and incurring high…
Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts
Luyang Fang, Tao Wang, Ping Ma +1
Automated scoring of written constructed responses typically relies on separate models per task, straining computational resources, storage, and maintenance in real-world education…
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
Haoran Lu, Luyang Fang, Ruidong Zhang +47
Due to the remarkable capabilities and growing impact of large language models (LLMs), they have been deeply integrated into many aspects of society. Thus, ensuring their alignment…
Non-Destructive Peat Analysis using Hyperspectral Imaging and Machine Learning
Yijun Yan, Jinchang Ren, Barry Harrison +3
Peat, a crucial component in whisky production, imparts distinctive and irreplaceable flavours to the final product. However, the extraction of peat disrupts ancient ecosystems and…
S2MNet: Speckle-To-Mesh Net for Three-Dimensional Cardiac Morphology Reconstruction via Echocardiogram
Xilin Gong, Yongkai Chen, Shushan Wu +3
Echocardiogram is the most commonly used imaging modality in cardiac assessment duo to its non-invasive nature, real-time capability, and cost-effectiveness. Despite its advantages…
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
Ehsan Latif, Luyang Fang, Ping Ma +1
This study proposes a method for knowledge distillation (KD) of fine-tuned Large Language Models (LLMs) into smaller, more efficient, and accurate neural networks. We specifically…
Online Sequential Leveraging Sampling Method for Streaming Autoregressive Time Series with Application to Seismic Data
Rui Xie, T. N. Sriram, Wei Biao Wu +1
Seismic data contain complex temporal information that arrives at high speed and has a large, even potentially unbounded volume. The explosion of temporally correlated streaming da…
Optimal Penalized Function-on-Function Regression under a Reproducing Kernel Hilbert Space Framework
Xiaoxiao Sun, Pang Du, Xiao Wang +1
Many scientific studies collect data where the response and predictor variables are both functions of time, location, or some other covariate. Understanding the relationship betwee…
Optimal Subsampling Approaches for Large Sample Linear Regression
Rong Zhu, Ping Ma, Michael W. Mahoney +1
A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample…
Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation
Luyang Fang, Haoran Lu, Yongkai Chen +2
As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge distillation reduces deployment cos…
Leverage classifier: Another look at support vector machine
Yixin Han, Jun Yu, Nan Zhang +4
Support vector machine (SVM) is a popular classifier known for accuracy, flexibility, and robustness. However, its intensive computation has hindered its application to large-scale…
Wahkon: A Statistically Principled Deep RKHS Superposition Network
Yongkai Chen, Wenxuan Zhong, Ping Ma
Deep learning excels at prediction but often lacks finite-sample guarantees and calibrated uncertainty; RKHS (Reproducing Kernel Hilbert Space)-based methods provide those guarante…
Semi-supervised Concept Bottleneck Models
Lijie Hu, Tianhao Huang, Huanyi Xie +6
Concept Bottleneck Models (CBMs) have garnered increasing attention due to their ability to provide concept-based explanations for black-box deep learning models while achieving hi…
Efficient Multi-Task Inferencing: Model Merging with Gromov-Wasserstein Feature Alignment
Luyang Fang, Ehsan Latif, Haoran Lu +3
Automatic scoring of student responses enhances efficiency in education, but deploying a separate neural network for each task increases storage demands, maintenance efforts, and r…
A Systematic Assessment of OpenAI o1-Preview for Higher Order Thinking in Education
Ehsan Latif, Yifan Zhou, Shuchen Guo +24
As artificial intelligence (AI) continues to advance, it demonstrates capabilities comparable to human intelligence, with significant potential to transform education and workforce…
LowCon: A design-based subsampling approach in a misspecified linear modeL
Cheng Meng, Rui Xie, Abhyuday Mandal +3
We consider a measurement constrained supervised learning problem, that is, (1) full sample of the predictors are given; (2) the response observations are unavailable and expensive…
NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning
Haoran Lu, Luyang Fang, Wenxuan Zhong +1
Multi-agent language systems are often built as hand-designed workflows, where agents are assigned semantic roles and communication protocols are specified in advance. We propose N…
Penalized Clustering of Large Scale Functional Data with Multiple Covariates
Ping Ma, Wenxuan Zhong
In this article, we propose a penalized clustering method for large scale data with multiple covariates through a functional data approach. In the proposed method, responses and co…
An optimal transport approach for selecting a representative subsample with application in efficient kernel density estimation
Jingyi Zhang, Cheng Meng, Jun Yu +3
Subsampling methods aim to select a subsample as a surrogate for the observed sample. Such methods have been used pervasively in large-scale data analytics, active learning, and pr…
Managing dataset shift by adversarial validation for credit scoring
Hongyi Qian, Baohui Wang, Ping Ma +3
Dataset shift is common in credit scoring scenarios, and the inconsistency between the distribution of training data and the data that actually needs to be predicted is likely to c…
Optimal Subsampling for Large Sample Logistic Regression
HaiYing Wang, Rong Zhu, Ping Ma
For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary…
Optimal smoothing in nonparametric mixed-effect models
Chong Gu, Ping Ma
Mixed-effect models are widely used for the analysis of correlated data such as longitudinal data and repeated measures. In this article, we study an approach to the nonparametric…