6 papers
ROAST: Rollout-based On-distribution Activation Steering Technique
Xuanbo Su, Hao Luo, Yingfang Zhang +1
Activation steering provides parameter-efficient control over large language models (LLMs) at inference time, but many methods rely on off-distribution supervision and discrete mas…
Parameter-Free Clustering via Self-Supervised Consensus Maximization (Extended Version)
Lijun Zhang, Suyuan Liu, Siwei Wang +4
Clustering is a fundamental task in unsupervised learning, but most existing methods heavily rely on hyperparameters such as the number of clusters or other sensitive settings, lim…
Convergence Analysis of the Lion Optimizer in Centralized and Distributed Settings
Wei Jiang, Lijun Zhang
In this paper, we analyze the convergence properties of the Lion optimizer. First, we establish that the Lion optimizer attains a convergence rate of …
Dual Adaptivity: Universal Algorithms for Minimizing the Adaptive Regret of Convex Functions
Lijun Zhang, Wenhao Yang, Guanghui Wang +2
To deal with changing environments, a new performance measure -- adaptive regret, defined as the maximum static regret over any interval, was proposed in online learning. Under the…
Improved Analysis for Sign-based Methods with Momentum Updates
Wei Jiang, Dingzhi Yu, Sifan Yang +2
In this paper, we present enhanced analysis for sign-based optimization algorithms with momentum updates. Traditional sign-based methods, under the separable smoothness assumption,…
Group Distributionally Robust Optimization with Flexible Sample Queries
Haomin Bai, Dingzhi Yu, Shuai Li +2
Group distributionally robust optimization (GDRO) aims to develop models that perform well across distributions simultaneously. Existing GDRO algorithms can only process a fixe…