3 papers
cs.CV2026
Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
Thomas Fel, Matthew Kowal, Mozes Jacobs +22
What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat concepts as isolated directio…
cs.LG2026
Symmetry Breaking in Transformers for Efficient and Interpretable Training
Eva Silverstein, Daniel Kunin, Vasudev Shyam
The attention mechanism in its standard implementation contains extraneous rotational degrees of freedom that are carried through computation but do not affect model activations or…
cs.LG2025
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
Vasudev Shyam, Jonathan Pilault, Emily Shepperd +2
Our formulation reveals that the reduction across the sequence axis can be efficiently computed in parallel through a tree reduction. Our algorithm, called Tree Attention, for para…