papers

Publications (87)

cs.LG2026

Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning in Agents

Neharika Jali, Anupam Nayak, Gauri Joshi

As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking traces even for simple queries. Prior appro…

cs.LG2026

Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning

Baris Askin, Holger R. Roth, Zhenyu Sun +3

Federated learning (FL) enables collaborative model training across distributed clients without sharing raw data, yet its scalability is limited by synchronization overhead. Asynch…

stat.ML2019

Correlated Multi-armed Bandits with a Latent Random Source

Samarth Gupta, Gauri Joshi, Osman Yağan

We consider a novel multi-armed bandit framework where the rewards obtained by pulling the arms are functions of a common latent random variable. The correlation between arms due t…

cs.LG2022

On the Unreasonable Effectiveness of Federated Averaging with Heterogeneous Data

Jianyu Wang, Rudrajit Das, Gauri Joshi +3

Existing theory predicts that data heterogeneity will degrade the performance of the Federated Averaging (FedAvg) algorithm in federated learning. However, in practice, the simple…

cs.LG2019

Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD

Jianyu Wang, Gauri Joshi

Large-scale machine learning training, in particular distributed stochastic gradient descent, needs to be robust to inherent system variability such as node straggling and random c…

cs.LG2025

Nonlinear Stochastic Gradient Descent and Heavy-tailed Noise: A Unified Framework and High-probability Guarantees

Aleksandar Armacki, Shuhua Yu, Pranay Sharma +4

We study high-probability convergence in online learning, in the presence of heavy-tailed noise. To combat the heavy tails, a general framework of nonlinear SGD methods is consider…

cs.LG2024

FedFisher: Leveraging Fisher Information for One-Shot Federated Learning

Divyansh Jhunjhunwala, Shiqiang Wang, Gauri Joshi

Standard federated learning (FL) algorithms typically require multiple rounds of communication between the server and the clients, which has several drawbacks, including requiring…

stat.ML2021

Best-Arm Identification in Correlated Multi-Armed Bandits

Samarth Gupta, Gauri Joshi, Osman Yağan

In this paper we consider the problem of best-arm identification in multi-armed bandits in the fixed confidence setting, where the goal is to identify, with probability for…

cs.LG2019

MATCHA: Speeding Up Decentralized SGD via Matching Decomposition Sampling

Jianyu Wang, Anit Kumar Sahu, Zhouyi Yang +2

This paper studies the problem of error-runtime trade-off, typically encountered in decentralized training based on stochastic gradient descent (SGD) using a given network. While a…

cs.LG2024

FedECADO: A Dynamical System Model of Federated Learning

Aayushya Agarwal, Gauri Joshi, Larry Pileggi

Federated learning harnesses the power of distributed optimization to train a unified machine learning model across separate clients. However, heterogeneous data distributions and…

cs.IT2014

Throughput-Smoothness Trade-offs in Multicasting of an Ordered Packet Stream

Gauri Joshi, Yuval Kochman, Gregory Wornell

An increasing number of streaming applications need packets to be strictly in-order at the receiver. This paper provides a framework for analyzing in-order packet delivery in such…

stat.ML2018

Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD

Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh +2

Distributed Stochastic Gradient Descent (SGD) when run in a synchronous manner, suffers from delays in waiting for the slowest learners (stragglers). Asynchronous methods can allev…

cs.LG2022

Heterogeneous Ensemble Knowledge Transfer for Training Large Models in Federated Learning

Yae Jee Cho, Andre Manoel, Gauri Joshi +2

Federated learning (FL) enables edge-devices to collaboratively learn a model without disclosing their private data to a central aggregating server. Most existing FL algorithms req…

cs.LG2026

MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models

Arian Raje, Anupam Nayak, Gauri Joshi

Mixture-of-Experts (MoE) model architectures can significantly reduce the number of activated parameters per token, enabling computationally efficient training and inference. Howev…

cs.LG2020

Bandit-based Communication-Efficient Client Selection Strategies for Federated Learning

Yae Jee Cho, Samarth Gupta, Gauri Joshi +1

Due to communication constraints and intermittent client availability in federated learning, only a subset of clients can participate in each training round. While most prior works…

cs.LG2025

Natural Policy Gradient for Average Reward Non-Stationary RL

Neharika Jali, Eshika Pathak, Pranay Sharma +2

We consider the problem of non-stationary reinforcement learning (RL) in the infinite-horizon average-reward setting. We model it by a Markov Decision Process with time-varying rew…

cs.LG2025

Adaptive Federated Learning via Dynamical System Model

Aayushya Agarwal, Larry Pileggi, Gauri Joshi

Hyperparameter selection is critical for stable and efficient convergence of heterogeneous federated learning, where clients differ in computational capabilities, and data distribu…

cs.LG2024

Erasure Coded Neural Network Inference via Fisher Averaging

Divyansh Jhunjhunwala, Neharika Jali, Gauri Joshi +1

Erasure-coded computing has been successfully used in cloud systems to reduce tail latency caused by factors such as straggling servers and heterogeneous traffic variations. A majo…

cs.LG2022

FedLite: A Scalable Approach for Federated Learning on Resource-constrained Clients

Jianyu Wang, Hang Qi, Ankit Singh Rawat +4

In classical federated learning, the clients contribute to the overall training by communicating local updates for the underlying model on their private data to a coordinating serv…

cs.LG2021

Advances and Open Problems in Federated Learning

Peter Kairouz, H. Brendan McMahan, Brendan Avent +56

Federated learning (FL) is a machine learning setting where many clients (e.g. mobile devices or whole organizations) collaboratively train a model under the orchestration of a cen…

cs.LG2026

Ravan: Multi-Head Low-Rank Adaptation for Federated Fine-Tuning

Arian Raje, Baris Askin, Divyansh Jhunjhunwala +1

Large language models (LLMs) have not yet effectively leveraged the vast amounts of edge-device data, and federated learning (FL) offers a promising paradigm to collaboratively fin…

cs.LG2020

Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization

Jianyu Wang, Qinghua Liu, Hao Liang +2

In federated optimization, heterogeneity in the clients' local datasets and computation speeds results in large variations in the number of local updates performed by each client i…

cs.IT2014

The Effect of Block-wise Feedback on the Throughput-Delay Trade-off in Streaming

Gauri Joshi, Yuval Kochman, Gregory Wornell

Unlike traditional file transfer where only total delay matters, streaming applications impose delay constraints on each packet and require them to be in order. To achieve fast in-…

cs.LG2020

Client Selection in Federated Learning: Convergence Analysis and Power-of-Choice Selection Strategies

Yae Jee Cho, Jianyu Wang, Gauri Joshi

Federated learning is a distributed optimization paradigm that enables a large number of resource-limited client nodes to cooperatively train a model without data sharing. Several…

stat.ML2021

Multi-Armed Bandits with Correlated Arms

Samarth Gupta, Shreyas Chaudhari, Gauri Joshi +1

We consider a multi-armed bandit framework where the rewards obtained by pulling different arms are correlated. We develop a unified approach to leverage these reward correlations…

cs.LG2019

Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms

Jianyu Wang, Gauri Joshi

Communication-efficient SGD algorithms, which allow nodes to perform local updates and periodically synchronize local models, are highly effective in improving the speed and scalab…

cs.LG2024

Debiasing Federated Learning with Correlated Client Participation

Zhenyu Sun, Ziyang Zhang, Zheng Xu +3

In cross-device federated learning (FL) with millions of mobile clients, only a small subset of clients participate in training in every communication round, and Federated Averagin…

cs.LG2021

Personalized Federated Learning for Heterogeneous Clients with Clustered Knowledge Transfer

Yae Jee Cho, Jianyu Wang, Tarun Chiruvolu +1

Personalized federated learning (FL) aims to train model(s) that can perform well for individual clients that are highly data and system heterogeneous. Most work in personalized FL…

cs.LG2021

Deep Kernels with Probabilistic Embeddings for Small-Data Learning

Ankur Mallick, Chaitanya Dwivedi, Bhavya Kailkhura +2

Gaussian Processes (GPs) are known to provide accurate predictions and uncertainty estimates even with small amounts of labeled data by capturing similarity between data points thr…

stat.ML2021

A Unified Approach to Translate Classical Bandit Algorithms to the Structured Bandit Setting

Samarth Gupta, Shreyas Chaudhari, Subhojyoti Mukherjee +2

We consider a finite-armed structured bandit problem in which mean rewards of different arms are known functions of a common hidden parameter . Since we do not place any rest…

cs.DC2014

Efficient Task Replication for Fast Response Times in Parallel Computation

Da Wang, Gauri Joshi, Gregory Wornell

One typical use case of large-scale distributed computing in data centers is to decompose a computation job into many independent tasks and run them in parallel on different machin…

cs.LG2026

Federate the Router: Learning Language Model Routers with Sparse and Decentralized Evaluations

Baris Askin, Shivam Patel, Anupam Nayak +4

Large language models (LLMs) are increasingly accessed as remotely hosted services by edge and enterprise clients that cannot run frontier models locally. Since models vary widely…

cs.LG2025

ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers

Shivam Patel, Neharika Jali, Ankur Mallick +1

Large language model (LLM) query routers are critical to modern AI platforms as they seek to improve efficiency by assigning inference queries to accurate, yet low-cost models. Par…

cs.LG2024

Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage Suffices

Jiin Woo, Laixi Shi, Gauri Joshi +1

Offline reinforcement learning (RL), which seeks to learn an optimal policy using offline data, has garnered significant interest due to its potential in critical applications wher…

cs.IT2012

Coding for Fast Content Download

Gauri Joshi, Yanpei Liu, Emina Soljanin

We study the fundamental trade-off between storage and content download time. We show that the download time can be significantly reduced by dividing the content into chunks, encod…

cs.LG2023

On the Convergence of Federated Averaging with Cyclic Client Participation

Yae Jee Cho, Pranay Sharma, Gauri Joshi +3

Federated Averaging (FedAvg) and its variants are the most popular optimization algorithms in federated learning (FL). Previous convergence analyses of FedAvg either assume full cl…

eess.SY2021

Job Dispatching Policies for Queueing Systems with Unknown Service Rates

Tuhinangshu Choudhury, Gauri Joshi, Weina Wang +1

In multi-server queueing systems where there is no central queue holding all incoming jobs, job dispatching policies are used to assign incoming jobs to the queue at one of the ser…

cs.LG2026

Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov Games

Anupam Nayak, Tong Yang, Osman Yagan +2

Reverse Kullback-Leibler (KL) divergence-based regularization with respect to a fixed reference policy is widely used in modern reinforcement learning to preserve the desired trait…

stat.ML2020

Slow and Stale Gradients Can Win the Race

Sanghamitra Dutta, Jianyu Wang, Gauri Joshi

Distributed Stochastic Gradient Descent (SGD) when run in a synchronous manner, suffers from delays in runtime as it waits for the slowest workers (stragglers). Asynchronous method…

cs.LG2024

Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation Models

Yae Jee Cho, Luyang Liu, Zheng Xu +2

Foundation models (FMs) adapt well to specific domains or tasks with fine-tuning, and federated learning (FL) enables the potential for privacy-preserving fine-tuning of the FMs wi…

cs.LG2024

Federated Stochastic Approximation under Markov Noise and Heterogeneity: Applications in Reinforcement Learning

Sajad Khodadadian, Pranay Sharma, Gauri Joshi +1

Since reinforcement learning algorithms are notoriously data-intensive, the task of sampling observations from the environment is usually split across multiple agents. However, tra…

cs.PF2022

Tackling Heterogeneous Traffic in Multi-access Systems via Erasure Coded Servers

Tuhinangshu Choudhury, Weina Wang, Gauri Joshi

Most data generated by modern applications is stored in the cloud, and there is an exponential growth in the volume of jobs to access these data and perform computations using them…

cs.LG2022

FedVARP: Tackling the Variance Due to Partial Client Participation in Federated Learning

Divyansh Jhunjhunwala, Pranay Sharma, Aushim Nagarkatti +1

Data-heterogeneous federated learning (FL) systems suffer from two significant sources of convergence error: 1) client drift error caused by performing multiple local optimization…

cs.LG2023

The Blessing of Heterogeneity in Federated Q-Learning: Linear Speedup and Beyond

Jiin Woo, Gauri Joshi, Yuejie Chi

When the data used for reinforcement learning (RL) are collected by multiple agents in a distributed manner, federated versions of RL algorithms allow collaborative learning withou…

cs.LG2024

Optimized Tradeoffs for Private Prediction with Majority Ensembling

Shuli Jiang, Qiuyi, Zhang +1

We study a classical problem in private prediction, the problem of computing an -differentially private majority of -differentially private algorithms for…

cs.DC2013

On the Delay-Storage Trade-off in Content Download from Coded Distributed Storage Systems

Gauri Joshi, Yanpei Liu, Emina Soljanin

In this paper we study how coding in distributed storage reduces expected download time, in addition to providing reliability against disk failures. The expected download time is r…

cs.LG2020

Machine Learning on Volatile Instances

Xiaoxi Zhang, Jianyu Wang, Gauri Joshi +1

Due to the massive size of the neural network models and training datasets used in machine learning today, it is imperative to distribute stochastic gradient descent (SGD) by split…

cs.LG2019

MLSys: The New Frontier of Machine Learning Systems

Alexander Ratner, Dan Alistarh, Gustavo Alonso +66

Machine learning (ML) techniques are enjoying rapidly increasing adoption. However, designing and implementing the systems that support ML models in real-world deployments remains…

cs.LG2020

Probabilistic Neighbourhood Component Analysis: Sample Efficient Uncertainty Estimation in Deep Learning

Ankur Mallick, Chaitanya Dwivedi, Bhavya Kailkhura +2

While Deep Neural Networks (DNNs) achieve state-of-the-art accuracy in various applications, they often fall short in accurately estimating their predictive uncertainty and, in tur…

cs.AI2026

Internal Planning in Language Models: Characterizing Horizon and Branch Awareness

Muhammed Ustaomeroglu, Baris Askin, Gauri Joshi +2

The extent to which decoder-only language models (LMs) engage in planning, that is, organizing intermediate computations to support coherent long-range generation, remains an impor…

cs.LG2025

Federated Communication-Efficient Multi-Objective Optimization

Baris Askin, Pranay Sharma, Gauri Joshi +1

We study a federated version of multi-objective optimization (MOO), where a single model is trained to optimize multiple objective functions. MOO has been extensively studied in th…

cs.DC2019

Rateless Codes for Near-Perfect Load Balancing in Distributed Matrix-Vector Multiplication

Ankur Mallick, Malhar Chaudhari, Utsav Sheth +2

Large-scale machine learning and data mining applications require computer systems to perform massive matrix-vector and matrix-matrix multiplication operations that need to be para…

cs.LG2018

Active Distribution Learning from Indirect Samples

Samarth Gupta, Gauri Joshi, Osman Yağan

This paper studies the problem of {\em learning} the probability distribution of a discrete random variable using indirect and sequential samples. At each time step, we c…

cs.DC2017

Efficient Straggler Replication in Large-scale Parallel Computing

Da Wang, Gauri Joshi, Gregory Wornell

In a cloud computing job with many parallel tasks, the tasks on the slowest machines (straggling tasks) become the bottleneck in the job completion. Computing frameworks such as Ma…

cs.LG2026

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR

Anupam Nayak, Baris Askin, Muhammed Ustaomeroglu +2

Reasoning post-training with reinforcement learning from verifiable rewards (RLVR) is typically studied in centralized settings, yet many realistic applications involve decentraliz…

cs.LG2024

Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems

Neharika Jali, Guannan Qu, Weina Wang +1

We consider the problem of efficiently routing jobs that arrive into a central queue to a system of heterogeneous servers. Unlike homogeneous systems, a threshold policy, that rout…

cs.LG2020

Overlap Local-SGD: An Algorithmic Approach to Hide Communication Delays in Distributed SGD

Jianyu Wang, Hao Liang, Gauri Joshi

Distributed stochastic gradient descent (SGD) is essential for scaling the machine learning algorithms to a large number of computing nodes. However, the infrastructures variabilit…

cs.IT2019

Service Rate Region of Content Access from Erasure Coded Storage

Sarah Anderson, Ann Johnston, Gauri Joshi +3

We consider storage systems in which files are stored over nodes. A node may be systematic for a particular file in the sense that access to it gives access to the file. Al…

cs.LG2023

Maximizing Global Model Appeal in Federated Learning

Yae Jee Cho, Divyansh Jhunjhunwala, Tian Li +2

Federated learning typically considers collaboratively training a global model using local data at edge clients. Clients may have their own individual requirements, such as having…

cs.LG2021

Adaptive Quantization of Model Updates for Communication-Efficient Federated Learning

Divyansh Jhunjhunwala, Advait Gadhikar, Gauri Joshi +1

Communication of model updates between client nodes and the central aggregating server is a major bottleneck in federated learning, especially in bandwidth-limited settings and hig…

cs.LG2021

Local Adaptivity in Federated Learning: Convergence and Consistency

Jianyu Wang, Zheng Xu, Zachary Garrett +3

The federated learning (FL) framework trains a machine learning model using decentralized data stored at edge client devices by periodically aggregating locally trained models. Pop…

math.OC2022

Federated Minimax Optimization: Improved Convergence Analyses and Algorithms

Pranay Sharma, Rohan Panda, Gauri Joshi +1

In this paper, we consider nonconvex minimax optimization, which is gaining prominence in many modern machine learning applications such as GANs. Large-scale edge-based collection…

cs.LG2022

Multi-Model Federated Learning with Provable Guarantees

Neelkamal Bhuyan, Sharayu Moharir, Gauri Joshi

Federated Learning (FL) is a variant of distributed learning where edge devices collaborate to learn a model without sharing their data with the central server or each other. We re…

cs.IT2015

On Throughput-Smoothness Trade-offs in Streaming Communication

Gauri Joshi, Yuval Kochman, Gregory Wornell

Unlike traditional file transfer where only total delay matters, streaming applications impose delay constraints on each packet and require them to be in order. To achieve fast in-…

cs.LG2024

High-probability Convergence Bounds for Nonlinear Stochastic Gradient Descent Under Heavy-tailed Noise

Aleksandar Armacki, Pranay Sharma, Gauri Joshi +3

We study high-probability convergence guarantees of learning on streaming data in the presence of heavy-tailed noise. In the proposed scenario, the model is updated in an online fa…

cs.DC2017

Efficient Redundancy Techniques for Latency Reduction in Cloud Systems

Gauri Joshi, Emina Soljanin, Gregory Wornell

In cloud computing systems, assigning a task to multiple servers and waiting for the earliest copy to finish is an effective method to combat the variability in response time of in…

cs.LG2026

LOCUS: Low-Dimensional Model Embeddings for Efficient Model Exploration, Comparison, and Selection

Shivam Patel, William Cocke, Gauri Joshi

The rapidly growing ecosystem of Large Language Models (LLMs) makes it increasingly challenging to manage and utilize the vast and dynamic pool of models effectively. We propose LO…

cs.LG2026

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

Baris Askin, Muhammed Ustaomeroglu, Anupam Nayak +3

Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that e…

cs.LG2021

A Field Guide to Federated Optimization

Jianyu Wang, Zachary Charles, Zheng Xu +50

Federated learning and analytics are a distributed approach for collaboratively learning models (or statistics) from decentralized data, motivated by and designed for privacy prote…

cs.LG2021

Leveraging Spatial and Temporal Correlations in Sparsified Mean Estimation

Divyansh Jhunjhunwala, Ankur Mallick, Advait Gadhikar +2

We study the problem of estimating at a central server the mean of a set of vectors distributed across several nodes (one vector per node). When the vectors are high-dimensional, t…

cs.LG2023

Federated Learning under Distributed Concept Drift

Ellango Jothimurugesan, Kevin Hsieh, Jianyu Wang +2

Federated Learning (FL) under distributed concept drift is a largely unexplored area. Although concept drift is itself a well-studied phenomenon, it poses particular challenges for…

cs.DC2020

Synergy via Redundancy: Adaptive Replication Strategies and Fundamental Limits

Gauri Joshi, Dhruva Kaushal

The maximum possible throughput (or the rate of job completion) of a multi-server system is typically the sum of the service rates of individual servers. Recent work shows that lau…

cs.DC2015

Efficient Replication of Queued Tasks for Latency Reduction in Cloud Systems

Gauri Joshi, Emina Soljanin, Gregory Wornell

In cloud computing systems, assigning a job to multiple servers and waiting for the earliest copy to finish is an effective method to combat the variability in response time of ind…

cs.AI2026

Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads

Shivam Patel, Akaash R. Parthasarathy, Ankur Mallick +1

Modern language query routers improve inference efficiency by assigning each query to a model that balances response quality and monetary cost. However, current query routers are l…

cs.LG2024

FedAST: Federated Asynchronous Simultaneous Training

Baris Askin, Pranay Sharma, Carlee Joe-Wong +1

Federated Learning (FL) enables edge devices or clients to collaboratively train machine learning (ML) models without sharing their private data. Much of the existing work in FL fo…

cs.IT2021

Service Rate Region: A New Aspect of Coded Distributed System Design

Mehmet Aktas, Gauri Joshi, Swanand Kadhe +2

Erasure coding has been recognized as a powerful method to mitigate delays due to slow or straggling nodes in distributed systems. This work shows that erasure coding of data objec…

cs.IT2017

On the Service Capacity Region of Accessing Erasure Coded Content

Mehmet Aktas, Sarah E. Anderson, Ann Johnston +5

Cloud storage systems generally add redundancy in storing content files such that files are replicated or erasure coded and stored on nodes. In addition to providing re…

cs.LG2025

FedRPCA: Enhancing Federated LoRA Aggregation Using Robust PCA

Divyansh Jhunjhunwala, Arian Raje, Madan Ravi Ganesh +6

LoRA has emerged as one of the most promising fine-tuning techniques, especially for federated learning (FL), since it significantly reduces communication and computation costs at…

cs.CV2026

Navigating the Accuracy-Size Trade-Off with Flexible Model Merging

Akash Dhasade, Divyansh Jhunjhunwala, Milos Vujasinovic +2

Model merging has emerged as an efficient method to combine multiple single-task fine-tuned models. The merged model can enjoy multi-task capabilities without expensive training. W…

cs.LG2026

Improving the Convergence of Private Shuffled Gradient Methods with Public Data

Shuli Jiang, Pranay Sharma, Zhiwei Steven Wu +1

We consider the problem of differentially private (DP) convex empirical risk minimization (ERM). While the standard DP-SGD algorithm is theoretically well-established, practical im…

cs.DC2023

Correlation Aware Sparsified Mean Estimation Using Random Projection

Shuli Jiang, Pranay Sharma, Gauri Joshi

We study the problem of communication-efficient distributed vector mean estimation, a commonly used subroutine in distributed optimization and Federated Learning (FL). Rand- spa…

cs.LG2023

Federated Minimax Optimization with Client Heterogeneity

Pranay Sharma, Rohan Panda, Gauri Joshi

Minimax optimization has seen a surge in interest with the advent of modern applications such as GANs, and it is inherently more challenging than simple minimization. The difficult…

stat.ML2026

Sample Complexity of Average-Reward Q-Learning: From Single-agent to Federated Reinforcement Learning

Yuchen Jiao, Jiin Woo, Gen Li +2

Average-reward reinforcement learning offers a principled framework for long-term decision-making by maximizing the mean reward per time step. Although Q-learning is a widely used…

cs.LG2019

Accelerating Deep Learning by Focusing on the Biggest Losers

Angela H. Jiang, Daniel L. -K. Wong, Giulio Zhou +8

This paper introduces Selective-Backprop, a technique that accelerates the training of deep neural networks (DNNs) by prioritizing examples with high loss at each iteration. Select…

cs.LG2023

Local or Global: Selective Knowledge Assimilation for Federated Learning with Limited Labels

Yae Jee Cho, Gauri Joshi, Dimitrios Dimitriadis

Many existing FL methods assume clients with fully-labeled data, while in realistic settings, clients have limited labels due to the expensive and laborious process of labeling. Li…

cs.LG2023

FedExP: Speeding Up Federated Averaging via Extrapolation

Divyansh Jhunjhunwala, Shiqiang Wang, Gauri Joshi

Federated Averaging (FedAvg) remains the most popular algorithm for Federated Learning (FL) optimization due to its simple implementation, stateless nature, and privacy guarantees…

cs.LG2025

Initialization Matters: Unraveling the Impact of Pre-Training on Federated Learning

Divyansh Jhunjhunwala, Pranay Sharma, Zheng Xu +1

Initializing with pre-trained models when learning on downstream tasks is becoming standard practice in machine learning. Several recent works explore the benefits of pre-trained i…