3 papers
cs.LG2026
How Many Different Outputs Can a Transformer Generate?
Maxime Meyer, Mario Michelessa, Caroline Chaux +1
We study how we can leverage only a handful of characteristics of a transformer's architecture to closely predict the number of different sequences it can output, both qualitativel…
cs.LG2026
Bandit Convex Optimization with Gradient Prediction Adaptivity
Shuche Wang, Adarsh Barik, Vincent Y. F. Tan
Bandit convex optimization (BCO) is a fundamental online learning framework with partial feedback, where the learner observes only the loss incurred at the chosen decision point in…
eess.SP2025
A Mirror Descent-Based Algorithm for Corruption-Tolerant Distributed Gradient Descent
Shuche Wang, Vincent Y. F. Tan
Distributed gradient descent algorithms have come to the fore in modern machine learning, especially in parallelizing the handling of large datasets that are distributed across sev…