26 citations · 31 across the 16 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
Dake Bu, Wei Huang, Andi Han +4
Transformer-based large language models (LLMs) have displayed remarkable creative prowess and emergence capabilities. Existing empirical studies have revealed a strong connection b…
cs.LG2024
Improved Particle Approximation Error for Mean Field Neural Networks
Atsushi Nitanda
Mean-field Langevin dynamics (MFLD) minimizes an entropy-regularized nonlinear convex functional defined over the space of probability distributions. MFLD has gained attention due…