3 papers
cs.LG2026
The price of multi-group transductive learning
Noah Bergam, Samuel Deng, Daniel Hsu
We show every multi-group learner in the transductive setting may incur a multiplicative penalty in its error rate on some group relative to the error rate achievable in the single…
cs.LG2026
Fixed Universal Transformers
Jingwen Liu, Alexandr Andoni, Daniel Hsu
We introduce \emph{universal transformers}: fixed transformers that can simulate any transformer in a given class via a suitable input embedding. Analogous to a universal Turing ma…
cs.LG2025
Fast attention mechanisms: a tale of parallelism
Jingwen Liu, Hantao Yu, Clayton Sanford +2
Transformers have the representational capacity to simulate Massively Parallel Computation (MPC) algorithms, but they suffer from quadratic time complexity, which severely limits t…