2 papers
cs.LG2026
Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View
Gyuryang Heo, Timothy Ngotiaoco, Kazuki Irie +2
Scalable sequence models, such as Transformer variants and structured state-space models, often trade expressivity power for sequence-level parallelism, which enables efficient tra…
cs.LG2024
Do Mice Grok? Glimpses of Hidden Progress During Overtraining in Sensory Cortex
Tanishq Kumar, Blake Bordelon, Cengiz Pehlevan +2
Does learning of task-relevant representations stop when behavior stops changing? Motivated by recent theoretical advances in machine learning and the intuitive observation that hu…