4 papers
Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams
Siddharth Pal, Viktoria Rojkova
Context engineering decides what information a model carries forward, and current designs meter it in tokens: compressing the past into a bounded recurrent state, keeping a key-val…
Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention
Siddharth Pal, Viktoria Rojkova
Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keepin…
A Stationarity-and-Coupling Criterion for Training-Free Time-Lagged Spectral Embeddings of Multivariate Time Series
Siddharth Pal, Viktoria Rojkova
We study training-free fixed-length descriptors for multivariate time series and ask not merely whether such a descriptor performs well, but when it can be expected to work at all.…
Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior
Vyom Agarwal, Mokshda Gangrade, Siddharth Pal +1
Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compressio…