38 citations · 53 across the 10 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Predicting Emergent Capabilities by Finetuning
Charlie Snell, Eric Wallace, Dan Klein +1
A fundamental open challenge in modern LLM scaling is the lack of understanding around emergent capabilities. In particular, language model pretraining loss is known to be highly p…
cs.LG2012★ 6 cited
Mixture-of-Parents Maximum Entropy Markov Models
David S. Rosenberg, Dan Klein, Ben Taskar
We present the mixture-of-parents maximum entropy Markov model (MoP-MEMM), a class of directed graphical models extending MEMMs. The MoP-MEMM allows tractable incorporation of long…