4 papers
Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
Thomas Fel, Matthew Kowal, Mozes Jacobs +22
What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat concepts as isolated directio…
Symmetry Breaking in Transformers for Efficient and Interpretable Training
Eva Silverstein, Daniel Kunin, Vasudev Shyam
The attention mechanism in its standard implementation contains extraneous rotational degrees of freedom that are carried through computation but do not affect model activations or…
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
Vasudev Shyam, Jonathan Pilault, Emily Shepperd +2
Our formulation reveals that the reduction across the sequence axis can be efficiently computed in parallel through a tree reduction. Our algorithm, called Tree Attention, for para…
The Zamba2 Suite: Technical Report
Paolo Glorioso, Quentin Anthony, Yury Tokpanov +5
In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance again…