activity
20192024
most citedPaLM: Scaling Language Modeling with Pathways

2.1k citations · 6.8k across the 40 of their papers we have counts for

collaborators
Showing cs.LGShow all

16 papers · 1 filter

cs.LG2024

Deep State-Space Generative Model For Correlated Time-to-Event Predictions

Yuan Xue, Denny Zhou, Nan Du +4

Capturing the inter-dependencies among multiple types of clinically-critical events is critical not only to accurate future event prediction, but also to better treatment planning.…

cs.LG2023★ 96 cited

Large Language Models as Optimizers

Chengrun Yang, Xuezhi Wang, Yifeng Lu +4

Optimization is ubiquitous. While derivative-based algorithms have been powerful tools for various problems, the absence of gradient imposes challenges on many real-world applicati…

cs.LG2023★ 4 cited

Not All Semantics are Created Equal: Contrastive Self-supervised Learning with Automatic Temperature Individualization

Zi-Hao Qiu, Quanqi Hu, Zhuoning Yuan +3

In this paper, we aim to optimize a contrastive loss with individualized temperatures in a principled and systematic manner for self-supervised learning. The common practice of usi…

cs.LG2023★ 20 cited

Large Language Models as Tool Makers

Tianle Cai, Xuezhi Wang, Tengyu Ma +2

Recent research has highlighted the potential of large language models (LLMs) to improve their problem-solving capabilities with the aid of suitable external tools. In our work, we…

cs.LG2022★ 86 cited

What learning algorithm is in-context learning? Investigations with linear models

Ekin Akyürek, Dale Schuurmans, Jacob Andreas +2

Neural sequence models, especially transformers, exhibit a remarkable capacity for in-context learning. They can construct new predictors from sequences of labeled examples $(x, f(…

cs.LG2022★ 1.2k cited

Scaling Instruction-Finetuned Language Models

Hyung Won Chung, Le Hou, Shayne Longpre +32

Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we expl…