2 papers
cs.NE2024
Self-Assembly of a Biologically Plausible Learning Circuit
Qianli Liao, Liu Ziyin, Yulu Gan +3
Over the last four decades, the amazing success of deep learning has been driven by the use of Stochastic Gradient Descent (SGD) as the main optimization technique. The default imp…
cs.CL2024
On the Power of Decision Trees in Auto-Regressive Language Modeling
Yulu Gan, Tomer Galanti, Tomaso Poggio +1
Originally proposed for handling time series data, Auto-regressive Decision Trees (ARDTs) have not yet been explored for language modeling. This paper delves into both the theoreti…