Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Switch EMA: A Free Lunch for Better Flatness and Sharpness
Siyuan Li, Zicheng Liu, Juanxi Tian +9
Exponential Moving Average (EMA) is a widely used weight averaging (WA) regularization to learn flat optima for better generalizations without extra cost in deep neural network (DN…
cs.LG2024
Dataset Growth
Ziheng Qin, Zhaopan Xu, Yukun Zhou +10
Deep learning benefits from the growing abundance of available data. Meanwhile, efficiently dealing with the growing data scale has become a challenge. Data publicly available are…