2 papers
cs.CV2026
ChannelTok: Efficient Flexible-Length Vision Tokenization
Sukriti Paul, Arpit Bansal, Tom Goldstein
Leading flexible vision tokenizers achieve SOTA quality at an extreme cost, relying on parameter-heavy backbones and slow, multi-step generative decoders. We depart from this compl…
cs.LG2024
Transformers Can Do Arithmetic with the Right Embeddings
Sean McLeish, Arpit Bansal, Alex Stein +8
The poor performance of transformers on arithmetic tasks seems to stem in large part from their inability to keep track of the exact position of each digit inside of a large span o…