9 papers
Will Scaling Improve Social Simulation with LLMs?
Caleb Ziems, William Held, Su Doga Karaca +3
Large Language Model (LLM) social simulations are a promising research method, but they are not yet faithful enough to be adopted widely. In this work, we investigate whether the c…
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
Woody Haosheng Gan, William Held, Diyi Yang
The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate wh…
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
Potsawee Manakul, Woody Haosheng Gan, Martijn Bartelds +3
Current audio language models are predominantly text-first, either extending pre-trained text LLM backbones or relying on semantic-only audio tokens, limiting general audio modelin…
Relative Scaling Laws for LLMs
William Held, David Hall, Percy Liang +1
Scaling laws describe how language models improve with additional data, parameters, and compute. While widely used, they are typically measured on aggregate test sets. Aggregate ev…
Culture Cartography: Mapping the Landscape of Cultural Knowledge
Caleb Ziems, William Held, Jane Yu +3
To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training. How do we find such knowledge that is (1) salient…
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
Potsawee Manakul, Woody Haosheng Gan, Michael J. Ryan +5
Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor corre…