activity
20122026
most citedFinite Time Analysis of Linear Two-timescale Stochastic Approximation with Markovian Noise

26 citations · 117 across the 52 of their papers we have counts for

collaborators
Showing cs.LGShow all

13 papers · 1 filter

cs.LG2026

EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization

Chung-Yiu Yau, Dawei Li, Athanasios Glentis +3

Lookahead-based acceleration methods, such as Nesterov's momentum, are widely used in optimization, but they often become unreliable in deep learning training mainly due to stochas…

cs.LG2025

Learning Graph from Smooth Signals under Partial Observation: A Robustness Analysis

Hoang-Son Nguyen, Hoi-To Wai

Learning the graph underlying a networked system from nodal signals is crucial to downstream tasks in graph signal processing and machine learning. The presence of hidden nodes who…

cs.LG2025

Stochastic Gradient Descent with Strategic Querying

Nanfei Jiang, Hoi-To Wai, Mahnoosh Alizadeh

This paper considers a finite-sum optimization problem under first-order queries and investigates the benefits of strategic querying on stochastic gradient-based methods compared t…

cs.LG2025

Federated Majorize-Minimization: Beyond Parameter Aggregation

Aymeric Dieuleveut, Gersende Fort, Mahmoud Hegazy +1

This paper proposes a unified approach for designing stochastic optimization algorithms that robustly scale to the federated learning setting. Our work studies a class of Majorize-…

cs.LG2025

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models

Quan Wei, Chung-Yiu Yau, Hoi-To Wai +4

Supervised fine-tuning is a standard method for adapting pre-trained large language models (LLMs) to downstream tasks. Quantization has been recently studied as a post-training tec…

cs.LG2025

Multilinear Tensor Low-Rank Approximation for Policy-Gradient Methods in Reinforcement Learning

Sergio Rozada, Hoi-To Wai, Antonio G. Marques

Reinforcement learning (RL) aims to estimate the action to take given a (time-varying) state, with the goal of maximizing a cumulative reward function. Predominantly, there are two…