activity
20142025
most citedLarger language models do in-context learning differently

100 citations · 245 across the 16 of their papers we have counts for

collaborators

13 papers

cs.LG20241 cited

Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Kaiyue Wen, Zhiyuan Li, Jason Wang +3

Training language models currently requires pre-determining a fixed compute budget because the typical cosine learning rate schedule depends on the total number of steps. In contra…

cs.LG20241 cited

Formal Theorem Proving by Rewarding LLMs to Decompose Proofs Hierarchically

Kefan Dong, Arvind Mahankali, Tengyu Ma

Mathematical theorem proving is an important testbed for large language models' deep and abstract reasoning capability. This paper focuses on improving LLMs' ability to write proof…

cs.CV2023

Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation

Xiaohan Cui, Long Ma, Tengyu Ma +3

Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for re…

cs.LG20232 cited

Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better Generalization

Kaiyue Wen, Zhiyuan Li, Tengyu Ma

Despite extensive studies, the underlying reason as to why overparameterized neural networks can generalize remains elusive. Existing theory shows that common stochastic optimizers…

cs.LG2023

Toward -recovery of Nonlinear Functions: A Polynomial Sample Complexity Bound for Gaussian Random Fields

Kefan Dong, Tengyu Ma

Many machine learning applications require learning a function with a small worst-case error over the entire input domain, that is, the -error, whereas most existing theo…

cs.CL2023100 cited

Larger language models do in-context learning differently

Jerry Wei, Jason Wei, Yi Tay +8

We study how in-context learning (ICL) in language models is affected by semantic priors versus input-label mappings. We investigate two setups-ICL with flipped labels and ICL with…