Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
GRPOformer: Advancing Hyperparameter Optimization via Group Relative Policy Optimization
Haoxin Guo, Jiawen Pan, Weixin Zhai
Hyperparameter optimization (HPO) plays a critical role in improving model performance. Transformer-based HPO methods have shown great potential; however, existing approaches rely…
cs.LG2024
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
Ahan Gupta, Hao Guo, Yueming Yuan +2
Many efficient self-attention techniques have become prevalent since the inception of the transformer architecture. Two popular classes of these techniques a…
cs.LG2024
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Mateusz Åajszczak, Guillermo Cámbara, Yang Li +16
We introduce a text-to-speech (TTS) model called BASE TTS, which stands for ig daptive treamable TTS with mergent abilities. BASE TT…