Publications (91)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
Antonio Orvieto, Lin Xiao
We consider the problem of minimizing the average of a large number of smooth but possibly non-convex functions. In the context of most machine learning applications, each loss fun…
Salient Object Detection in Traffic Scene through the TSOD10K Dataset
Yu Qiu, Yuhang Sun, Jie Mei +2
Traffic Salient Object Detection (TSOD) aims to segment the objects critical to driving safety by combining semantic (e.g., collision risks) and visual saliency. Unlike SOD in natu…
DSCOVR: Randomized Primal-Dual Block Coordinate Algorithms for Asynchronous Distributed Optimization
Lin Xiao, Adams Wei Yu, Qihang Lin +1
Machine learning with big data often involves large optimization models. For distributed optimization over a cluster of machines, frequent communication and synchronization of all…
Using Statistics to Automate Stochastic Optimization
Hunter Lang, Pengchuan Zhang, Lin Xiao
Despite the development of numerous adaptive optimizers, tuning the learning rate of stochastic gradient methods remains a major roadblock to obtaining good practical performance i…
Secrecy Energy Efficiency Maximization for UAV-Enabled Mobile Relaying
Lin Xiao, Yu Xu, Dingcheng Yang +1
This paper investigates the secrecy energy efficiency (SEE) maximization problem for unmanned aerial vehicle enabled mobile relaying system, where a high-mobility UAV is exploited…
A Superluminous Supernova Lightened by Collisions with Pulsational Pair-instability Shells
Weili Lin, Xiaofeng Wang, Lin Yan +23
Superluminous supernovae are among the most energetic stellar explosions in the Universe, but their energy sources remain an open question. Here we present long-term observations o…
The environmental dependence of mid-IR luminous dusty Supernovae
Lin Xiao, Zeyue Peng, Lluis Galbany +10
Using the Spitzer and WISE images, we discovered 42 mid-IR luminous dusty supernovae with local integral-field spectroscopy data. The observed mid-IR emission indicates the presenc…
Very Late-Time JWST and Keck Spectra of the Oxygen-Rich Supernova 1995N
Geoffrey C. Clayton, R. Wesson, Ori D. Fox +43
We present new {\it JWST}/MIRI MRS and Keck spectra of SN 1995N obtained in 2022--2023, more than 10,000 days after the supernova (SN) explosion. These spectra are among the latest…
FedShuffle: Recipes for Better Use of Local Work in Federated Learning
Samuel Horváth, Maziar Sanjabi, Lin Xiao +2
The practice of applying several local updates before aggregation across clients has been empirically shown to be a successful approach to overcoming the communication bottleneck i…
The environmental dependence of Spitzer dusty Supernovae
Lin Xiao, Tamás Szalai, LluÃs Galbany +9
Thanks to the mid-infrared capability offered by Spitzer, systematic searches of dust in SNe have been carried out over the past decade. Studies have revealed the presence of a sub…
On Continual Model Refinement in Out-of-Distribution Data Streams
Bill Yuchen Lin, Sida Wang, Xi Victoria Lin +4
Real-world natural language processing (NLP) models need to be continually updated to fix the prediction errors in out-of-distribution (OOD) data streams while overcoming catastrop…
Spatially resolved MaNGA observations of the host galaxy of superluminous supernova 2017egm
Ting-Wan Chen, Patricia Schady, Lin Xiao +6
Superluminous supernovae (SLSNe) are found predominantly in dwarf galaxies, indicating that their progenitors have a low metallicity. However, the most nearby SLSN to date, SN 2017…
On the Complexity Analysis of Randomized Block-Coordinate Descent Methods
Zhaosong Lu, Lin Xiao
In this paper we analyze the randomized block-coordinate descent (RBCD) methods proposed in [8,11] for minimizing the sum of a smooth convex function and a block-separable convex f…
JWST Discovery of Dust Reservoirs in Nearby Type IIP Supernovae 2004et and 2017eaw
Melissa Shahbandeh, Arkaprabha Sarangi, Tea Temim +36
Supernova (SN) explosions have been sought for decades as a possible source of dust in the Universe, providing the seeds of galaxies, stars, and planetary systems. SN 1987A offers…
Multi-Level Composite Stochastic Optimization via Nested Variance Reduction
Junyu Zhang, Lin Xiao
We consider multi-level composite optimization problems where each mapping in the composition is the expectation over a family of random smooth mappings or the sum of some finite n…
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
Avinandan Bose, Zhihan Xiong, Yuejie Chi +3
Personalizing large language models (LLMs) to accommodate diverse user preferences is essential for enhancing alignment and user satisfaction. Traditional reinforcement learning fr…
The radius variations of accreting main sequence stars and mass transfer instability
Zi-Qi Zhao, Zhen-Wei Li, Lin Xiao +2
Many previous works studied the dynamical timescale mass transfer stability criteria based on the donor response with neglecting the stellar structure of the accretor. In this lett…
Label-aware Document Representation via Hybrid Attention for Extreme Multi-Label Text Classification
Xin Huang, Boli Chen, Lin Xiao +1
Extreme multi-label text classification (XMTC) aims at tagging a document with most relevant labels from an extremely large-scale label set. It is a challenging problem especially…
Robust Distributed Online Prediction
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir +1
The standard model of online prediction deals with serial processing of inputs by a single processor. However, in large-scale online prediction problems, where inputs arrive at a h…
Bregman Douglas-Rachford Splitting Method
Shiqian Ma, Lin Xiao, Renbo Zhao
In this paper, we propose the Bregman Douglas-Rachford splitting (BDRS) method and its variant Bregman Peaceman-Rachford splitting method for solving maximal monotone inclusion pro…
Fastest mixing Markov chain on graphs with symmetries
Stephen Boyd, Persi Diaconis, Pablo A. Parrilo +1
We show how to exploit symmetries of a graph to efficiently compute the fastest mixing Markov chain on the graph (i.e., find the transition probabilities on the edges to minimize t…
PARQ: Piecewise-Affine Regularized Quantization
Lisa Jin, Jianhao Ma, Zechun Liu +3
We develop a principled method for quantization-aware training (QAT) of large-scale machine learning models. Specifically, we show that convex, piecewise-affine regularization (PAR…
A Proximal Stochastic Gradient Method with Progressive Variance Reduction
Lin Xiao, Tong Zhang
We consider the problem of minimizing the sum of two convex functions: one is the average of a large number of smooth component functions, and the other is a general convex functio…
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
Tong Yang, Bo Dai, Lin Xiao +1
Multi-agent reinforcement learning (MARL) lies at the heart of a plethora of applications involving the interaction of a group of agents in a shared unknown environment. A prominen…
Statistical Adaptive Stochastic Gradient Methods
Pengchuan Zhang, Hunter Lang, Qiang Liu +1
We propose a statistical adaptive procedure called SALSA for automatically scheduling the learning rate (step size) in stochastic gradient methods. SALSA first uses a smoothed stoc…
From low probability to high confidence in stochastic convex optimization
Damek Davis, Dmitriy Drusvyatskiy, Lin Xiao +1
Standard results in stochastic convex optimization bound the number of samples that an algorithm needs to generate a point with small function value in expectation. More nuanced hi…
Core-Collapse Supernova Rate Synthesis Within 11 Mpc
Lin Xiao, J. J. Eldridge
The 11 Mpc H-alpha and Ultraviolet Galaxy (11HUGS) Survey traces the star formation activity of nearby galaxies. In addition within this volume the detection completeness of core-c…
SBEED: Convergent Reinforcement Learning with Nonlinear Function Approximation
Bo Dai, Albert Shaw, Lihong Li +5
When function approximation is used, solving the Bellman optimality equation with stability guarantees has remained a major open problem in reinforcement learning for decades. The…
Variational Gram Functions: Convex Analysis and Optimization
Amin Jalali, Maryam Fazel, Lin Xiao
We propose a new class of convex penalty functions, called \emph{variational Gram functions} (VGFs), that can promote pairwise relations, such as orthogonality, among a set of vect…
Noisy recovery from random linear observations: Sharp minimax rates under elliptical constraints
Reese Pathak, Martin J. Wainwright, Lin Xiao
Estimation problems with constrained parameter spaces arise in various settings. In many of these problems, the observations available to the statistician can be modelled as arisin…
Adaptive Stochastic Variance Reduction for Subsampled Newton Method with Cubic Regularization
Junyu Zhang, Lin Xiao, Shuzhong Zhang
The cubic regularized Newton method of Nesterov and Polyak has become increasingly popular for non-convex optimization because of its capability of finding an approximate local sol…
Accelerated Bregman Proximal Gradient Methods for Relatively Smooth Convex Optimization
Filip Hanzely, Peter Richtarik, Lin Xiao
We consider the problem of minimizing the sum of two convex functions: one is differentiable and relatively smooth with respect to a reference convex function, and the other can be…
Understanding the Role of Momentum in Stochastic Gradient Methods
Igor Gitman, Hunter Lang, Pengchuan Zhang +1
The use of momentum in stochastic gradient methods has become a widespread practice in machine learning. Different variants of momentum, including heavy-ball momentum, Nesterov's a…
Statistically Preconditioned Accelerated Gradient Method for Distributed Optimization
Hadrien Hendrikx, Lin Xiao, Sebastien Bubeck +2
We consider the setting of distributed empirical risk minimization where multiple machines compute the gradients in parallel and a centralized server updates the model parameters.…
Exploiting Strong Convexity from Data with Primal-Dual First-Order Algorithms
Jialei Wang, Lin Xiao
We consider empirical risk minimization of linear predictors with convex loss functions. Such problems can be reformulated as convex-concave saddle point problems, and thus are wel…
Faster Last-iterate Convergence of Policy Optimization in Zero-Sum Markov Games
Shicong Cen, Yuejie Chi, Simon S. Du +1
Multi-Agent Reinforcement Learning (MARL) -- where multiple agents learn to interact in a shared dynamic environment -- permeates across a wide range of critical applications. Whil…
On the Convergence Rates of Policy Gradient Methods
Lin Xiao
We consider infinite-horizon discounted Markov decision problems with finite state and action spaces and study the convergence rates of the projected policy gradient method and a g…
Joint Computation and Communication Design for UAV-Assisted Mobile Edge Computing in IoT
Tiankui Zhang, Yu Xu, Jonathan Loo +2
Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) system is a prominent concept, where a UAV equipped with a MEC server is deployed to serve a number of terminal d…
Serendipitous detection of the dusty Type IIL SN 1980K with JWST/MIRI
Szanna ZsÃros, Tamás Szalai, Ilse De Looze +39
We present mid-infrared (mid-IR) imaging of the Type IIL supernova (SN) 1980K with the James Webb Space Telescope (JWST) more than 40 yr post-explosion. SN 1980K, located in the ne…
A Randomized Nonmonotone Block Proximal Gradient Method for a Class of Structured Nonlinear Programming
Zhaosong Lu, Lin Xiao
We propose a randomized nonmonotone block proximal gradient (RNBPG) method for minimizing the sum of a smooth (possibly nonconvex) function and a block-separable (possibly nonconve…
Properties of the cores and filaments in the Ophiuchus molecular cloud and its L1688 hub-filament system
Bo-Sheng Jia, Guo-Yin Zhang, Alexander Menshchikov +9
Analyzing filaments and cores in molecular clouds is key to understanding galactic star formation and its environmental dependence. This paper studies the properties and distributi…
Does Head Label Help for Long-Tailed Multi-Label Text Classification
Lin Xiao, Xiangliang Zhang, Liping Jing +2
Multi-label text classification (MLTC) aims to annotate documents with the most relevant labels from a number of candidate labels. In real applications, the distribution of label f…
Stochastic optimization with decision-dependent distributions
Dmitriy Drusvyatskiy, Lin Xiao
Stochastic optimization problems often involve data distributions that change in reaction to the decision variables. This is the case for example when members of the population res…
Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies
Rui Yuan, Simon S. Du, Robert M. Gower +2
We consider infinite-horizon discounted Markov decision processes and study the convergence rates of the natural policy gradient (NPG) and the Q-NPG methods with the log-linear pol…
Stochastic Primal-Dual Coordinate Method for Regularized Empirical Risk Minimization
Yuchen Zhang, Lin Xiao
We consider a generic convex optimization problem associated with regularized empirical risk minimization of linear predictors. The problem structure allows us to reformulate it as…
An Accelerated Proximal Coordinate Gradient Method and its Application to Regularized Empirical Risk Minimization
Qihang Lin, Zhaosong Lu, Lin Xiao
We consider the problem of minimizing the sum of two convex functions: one is smooth and given by a gradient oracle, and the other is separable over blocks of coordinates and has a…
A Mid-infrared Study of Superluminous Supernovae
Luming Sun, Lin Xiao, Ge Li
We present the mid-infrared (MIR) light curves (LC) of 10 superluminous supernovae (SLSNe) at based on WISE data at 3.4 and 4.6 m. Three of them, including PS15br, SN…
Infrared Echoes of Optical Tidal Disruption Events: ~1% Dust Covering Factor or Less at sub-parsec Scale
Ning Jiang, Tinggui Wang, Xueyang Hu +3
The past decade has experienced an explosive increase of optically-discovered tidal disruption events (TDEs) with the advent of modern time-domain surveys. However, we still lack a…
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
Tong Yang, Bo Dai, Lin Xiao +1
Online reinforcement learning (RL) with complex function approximations such as transformers and deep neural networks plays a significant role in the modern practice of artificial…
Optimal Distributed Online Prediction using Mini-Batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir +1
Online prediction methods are typically presented as serial algorithms running on a single processor. However, in the age of web-scale prediction problems, it is increasingly commo…
Dual Approximation Policy Optimization
Zhihan Xiong, Maryam Fazel, Lin Xiao
We propose Dual Approximation Policy Optimization (DAPO), a framework that incorporates general function approximation into policy mirror descent methods. In contrast to the popula…
Improving Self-supervised Pre-training via a Fully-Explored Masked Language Model
Mingzhi Zheng, Dinghan Shen, Yelong Shen +2
Masked Language Model (MLM) framework has been widely adopted for self-supervised language pre-training. In this paper, we argue that randomly sampled masks in MLM would lead to un…
Prospects of Searching for Type Ia Supernovae with 2.5-m Wide Field Survey Telescope
Maokai Hu, Lei Hu, Ji-an Jiang +4
Type Ia Supernovae (SNe Ia) are the thermonuclear explosion of a carbon-oxygen white dwarf (WD) and are well-known as a distance indicator. However, it is still unclear how WDs inc…
Mass Ratio Distribution of Hierarchical Triple Systems from the LAMOST-MRS Survey
Tongyu He, Jiangdan Li, Xuefei Chen +3
Hierarchical triple-star systems consists of three components organised into an inner binary (,) and a more distant outer tertiary () star. The LAMOST Medium-R…
JWST/MIRI Observations of Newly Formed Dust in the Cold, Dense Shell of the Type IIn SN 2005ip
Melissa Shahbandeh, Ori D. Fox, Tea Temim +42
Dust from core-collapse supernovae (CCSNe), specifically Type IIP SNe, has been suggested to be a significant source of the dust observed in high-redshift galaxies. CCSNe eject lar…
Jointly Modeling Intra- and Inter-transaction Dependencies with Hierarchical Attentive Transaction Embeddings for Next-item Recommendation
Shoujin Wang, Longbing Cao, Liang Hu +4
A transaction-based recommender system (TBRS) aims to predict the next item by modeling dependencies in transactional data. Generally, two kinds of dependencies considered are intr…
Joint Linked Component Analysis for Multiview Data
Lin Xiao, Luo Xiao
In this work, we propose the joint linked component analysis (joint\_LCA) for multiview data. Unlike classic methods which extract the shared components in a sequential manner, the…
End-to-end Learning of LDA by Mirror-Descent Back Propagation over a Deep Architecture
Jianshu Chen, Ji He, Yelong Shen +5
We develop a fully discriminative learning approach for supervised Latent Dirichlet Allocation (LDA) model using Back Propagation (i.e., BP-sLDA), which maximizes the posterior pro…
Grad-GradaGrad? A Non-Monotone Adaptive Stochastic Gradient Method
Aaron Defazio, Baoyu Zhou, Lin Xiao
The classical AdaGrad method adapts the learning rate by dividing by the square root of a sum of squared gradients. Because this sum on the denominator is increasing, the method ca…
Communication-Efficient Distributed Optimization of Self-Concordant Empirical Loss
Yuchen Zhang, Lin Xiao
We consider distributed convex optimization problems originated from sample average approximation of stochastic optimization, or empirical risk minimization in machine learning. We…
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation
Tianxiang Xia, Lin Xiao, Yannick Montorfani +3
In this project, we address the issue of infidelity in text-to-image generation, particularly for actions involving multiple objects. For this we build on top of the CONFORM framew…
Seeded growth of high-quality transition metal dichalcogenide single crystals via chemical vapor transport
Hao Li, Junku Liu, Nan Guo +5
Transition metal dichalcogenides (TMDs) are van der Waals layered materials with sizable and tunable bandgaps, offering promising platforms for two-dimensional electronics and opto…
Sparse and Integrative Principal Component Analysis for Multiview Data
Lin Xiao, Luo Xiao
We consider dimension reduction of multiview data, which are emerging in scientific studies. Formulating multiview data as multi-variate data with block structures corresponding to…
Hyperbolic Interaction Model For Hierarchical Multi-Label Classification
Boli Chen, Xin Huang, Lin Xiao +2
Different from the traditional classification tasks which assume mutual exclusion of labels, hierarchical multi-label classification (HMLC) aims to assign multiple labels to every…
DiffCL: A Diffusion-Based Contrastive Learning Framework with Semantic Alignment for Multimodal Recommendations
Qiya Song, Jiajun Hu, Lin Xiao +3
Multimodal recommendation systems integrate diverse multimodal information into the feature representations of both items and users, thereby enabling a more comprehensive modeling…
Multimodal Graph Neural Network for Recommendation with Dynamic De-redundancy and Modality-Guided Feature De-noisy
Feng Mo, Lin Xiao, Qiya Song +2
Graph neural networks (GNNs) have become crucial in multimodal recommendation tasks because of their powerful ability to capture complex relationships between neighboring nodes. Ho…
Federated Learning with Partial Model Personalization
Krishna Pillutla, Kshitiz Malik, Abdelrahman Mohamed +3
We consider two federated learning algorithms for training partially personalized models, where the shared and personal parameters are updated either simultaneously or alternately…
A 3 mm line survey towards the circumstellar envelope of the carbon-rich AGB star IRC+10216 (CW Leo)
Juan Tuo, Xiaohu Li, Jixian Sun +14
We present an unbiased 3 mm spectral line survey (between 84.5 and 115.8 GHz), conducted by the Purple Mountain Observatory 13.7 meter radio telescope, together with updated m…
Cross-Platform Simulation Architecture with application to truck platooning impact assessment
Andres Ladino, Lin Xiao, Kingsley Adjenugwhure +2
Simulation-based traffic impact assessment studies of advanced technologies such as truck platooning need to be carried out to ascertain their benefits for traffic efficiency, safe…
The distance, supernova rate and supernova progenitors of NGC 6946
J. J. Eldridge, Lin Xiao
The distance to the fireworks galaxy NGC 6946 is highly uncertain. Recent distance estimates using the tip of the red giant branch of 7.7 to 7.8 Mpc are larger than the distance co…
Pairwise Instance Relation Augmentation for Long-tailed Multi-label Text Classification
Lin Xiao, Pengyu Xu, Liping Jing +1
Multi-label text classification (MLTC) is one of the key tasks in natural language processing. It aims to assign multiple target labels to one document. Due to the uneven popularit…
Emission-line Diagnostics of Nearby HII Regions Including Supernova Hosts
Lin Xiao, J. J. Eldridge, Elizabeth Stanway +1
We present a new model of the optical nebular emission from HII regions by combin- ing the results of the Binary Population and Spectral Synthesis (bpass) code with the photoion- i…
Chemical variations across the TMC-1 boundary: molecular tracers from translucent phase to dense phase
Long-Fei Chen, Di Li, Donghui Quan +4
We investigated the chemical evolutions of gas phase and grain surface species across the Taurus molecular cloud-1 (TMC-1) filament from translucent phase to dense phase. By compar…
A Stochastic Composite Gradient Method with Incremental Variance Reduction
Junyu Zhang, Lin Xiao
We consider the problem of minimizing the composition of a smooth (nonconvex) function and a smooth vector mapping, where the inner mapping is in the form of an expectation over so…
Core-collapse supernovae ages and metallicities from emission-line diagnostics of nearby stellar populations
Lin Xiao, L. Galbany, J. J. Eldridge +1
Massive stars are the main objects that illuminate H II regions and they evolve quickly to end their lives in core-collapse supernovae (CCSNe). Thus it is important to investigate…
JWST/MIRI detects the dusty SN1993J about 30 years after explosion
Tamás Szalai, Szanna ZsÃros, Jacob Jencson +37
Core-collapse supernovae (CCSNe) have long been considered to contribute significantly to the cosmic dust budget. New dust cools quickly and is therefore detectable at mid-infrared…
Stochastic Variance-Reduced Prox-Linear Algorithms for Nonconvex Composite Optimization
Junyu Zhang, Lin Xiao
We consider minimization of composite functions of the form , where and are convex functions (which can be nonsmooth) and is a smooth vector mapping. In a…
Learning SMaLL Predictors
Vikas K. Garg, Ofer Dekel, Lin Xiao
We present a new machine learning technique for training small resource-constrained predictors. Our algorithm, the Sparse Multiprototype Linear Learner (SMaLL), is inspired by the…
Online Classification Using a Voted RDA Method
Tianbing Xu, Jianfeng Gao, Lin Xiao +1
We propose a voted dual averaging method for online classification problems with explicit regularization. This method employs the update rule of the regularized dual averaging (RDA…
Importance Estimation from Multiple Perspectives for Keyphrase Extraction
Mingyang Song, Liping Jing, Lin Xiao
Keyphrase extraction is a fundamental task in Natural Language Processing, which usually contains two main parts: candidate keyphrase extraction and keyphrase importance estimation…
Direct Urca processes involving singlet proton superfluidity in neutron star cooling
Yan Xu, Xiu Lin Huang, Xiao Jun Zhang +4
A detailed description of the baryon direct Urca processes A: , B: , C: related to the…
Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
Jianhao Ma, Lin Xiao
Optimization problems over discrete or quantized variables are very challenging in general due to the combinatorial nature of their search space. Piecewise-affine regularization (P…
Stochastic Variance Reduction Methods for Policy Evaluation
Simon S. Du, Jianshu Chen, Lihong Li +2
Policy evaluation is a crucial step in many reinforcement-learning procedures, which estimates a value function that predicts states' long-term value under a given policy. In this…
Fast Clifford Neural Layers
Tianxiang Xia, Max Neuwinger, Lin Xiao
Clifford Neural Layers improve PDE modeling by introducing Clifford Algebra into neural networks. In this project we focus on optimizing the inference of 2/3D Clifford convolutiona…
Emission-line diagnostics of nearby HII regions including interacting binary populations
Lin Xiao, Elizabeth Stanway, J. J. Eldridge
We present numerical models of the nebular emission from H II regions around young stellar populations over a range of compositions and ages. The synthetic stellar pop- ulations in…
Cold-Start Personalization via Training-Free Priors from Structured World Models
Avinandan Bose, Shuyue Stella Li, Faeze Brahman +6
Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available. The core challenge is a routing problem: each…
A Proximal-Gradient Homotopy Method for the Sparse Least-Squares Problem
Lin Xiao, Tong Zhang
We consider solving the -regularized least-squares (-LS) problem in the context of sparse recovery, for applications such as compressed sensing. The standard proxim…
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
Aaron Defazio, Konstantin Mishchenko, Parameswaran Raman +2
We propose Generalized Primal Averaging (GPA), an extension of Nesterov's method that unifies and generalizes recent averaging-based optimizers like single-worker DiLoCo and Schedu…
BiT: Robustly Binarized Multi-distilled Transformer
Zechun Liu, Barlas Oguz, Aasish Pappu +5
Modern pre-trained transformers have rapidly advanced the state-of-the-art in machine learning, but have also grown in parameters and computational complexity, making them increasi…
Stochastic Approximation with Block Coordinate Optimal Stepsizes
Tao Jiang, Lin Xiao
We consider stochastic approximation with block-coordinate stepsizes and propose adaptive stepsize rules that aim to minimize the expected distance from the next iterate to an (unk…
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
Zechun Liu, Changsheng Zhao, Hanxian Huang +13
The optimal bit-width for achieving the best trade-off between quantized model size and accuracy has been a subject of ongoing debate. While some advocate for 4-bit quantization, o…