1 citations · 1 across the 2 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Induction Heads Interpolate N-Grams
Francesco D'Angelo, Oguz Kaan Yuksel, Swathi Shree Narashiman +1
Induction heads are attention circuits believed to underlie in-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive. We…
cs.IT2026
Semantic Smoothing for Language Models via Distribution Estimation and Embeddings
Haricharan Balasundaram, Swathi Shree Narashiman, Pranay Mathur +1
We propose semantic smoothing, a smoothing method for language models that uses embeddings to share statistical observations across semantically similar contexts. The starting poin…