3 papers
cs.LG2025
GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
Qifu Wen, Xi Zeng, Zihan Zhou +4
Early stopping monitors global validation loss and halts all parameter updates simultaneously, which is computationally costly for large transformers due to the extended time requi…
cs.LG2025
Discovering Hidden Algebraic Structures via Transformers with Rank-Aware Beam GRPO
Jaeha Lee, Gio Huh, Ning Su +1
Recent efforts have extended the capabilities of transformers in logical reasoning and symbolic computations. In this work, we investigate their capacity for non-linear latent patt…
cs.LG2024
Calibre: Towards Fair and Accurate Personalized Federated Learning with Self-Supervised Learning
Sijia Chen, Ningxin Su, Baochun Li
In the context of personalized federated learning, existing approaches train a global model to extract transferable representations, based on which any client could train personali…