6 papers
Sharp Capacity Thresholds in Linear Associative Memory: From Top-1 Retrieval to Tail-Average Learning
Nicholas Barnfield, Juno Kim, Eshaan Nichani +2
How many key-value associations can a linear memory store? The answer depends not only on the degrees of freedom in the memory matrix, but also on the retrieval c…
Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning
Nicholas Barnfield, Subhabrata Sen, Pragya Sur
Recent progress has rapidly advanced our understanding of the mechanisms underlying in-context learning in modern attention-based neural networks. However, existing results focus e…
High-Dimensional Analysis of Single-Layer Attention for Sparse-Token Classification
Nicholas Barnfield, Hugo Cui, Yue M. Lu
When and how can an attention mechanism learn to selectively attend to informative tokens, thereby enabling detection of weak, rare, and sparsely located features? We address these…
Ziv-Merhav estimation for hidden-Markov processes
Nicholas Barnfield, Raphaël Grondin, Gaia Pozzoli +1
We present a proof of strong consistency of a Ziv-Merhav-type estimator of the cross entropy rate for pairs of hidden-Markov processes. Our proof strategy has two novel aspects: th…
On the Ziv-Merhav theorem beyond Markovianity II: leveraging the thermodynamic formalism
Nicholas Barnfield, Raphaël Grondin, Gaia Pozzoli +1
We prove asymptotic results for a modification of the cross-entropy estimator originally introduced by Ziv and Merhav in the Markovian setting in 1993. Our results concern a more g…
On the Ziv-Merhav theorem beyond Markovianity
Nicholas Barnfield, Raphaël Grondin, Gaia Pozzoli +1
We generalize to a broader class of decoupled measures a result of Ziv and Merhav on universal estimation of the specific cross (or relative) entropy for a pair of multi-level Mark…