34 papers
Aligning Language Model Benchmarks with Pairwise Preferences
Marco Gutierrez, Xinyi Leng, Hannah Cyberey +3
Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real…
A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models
Soham Dan, Himanshu Beniwal, Thomas Hartvigsen
Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural contexts. This survey synthesizes w…
Inferring Events from Time Series using Language Models
Mingtian Tan, Mike A. Merrill, Zack Gottesman +3
A common goal in analyzing time series data is to understand how events cause observed variations. We study whether Large Language Models (LLMs) can infer natural language events a…
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
Xu Ouyang, Deyi Liu, Yuhang Cai +5
Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and q…
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
Natalie Seah, Danielle S. Bitterman, Daphna Spiegel +1
Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where clinicians must clearly and…
Test-Time Hinting for Black-Box Vision-Language Models
Kaihua Hou, Abhijith Varma Mudunuri, Jiaxing Qiu +3
Test-time scaling (TTS) methods have proven highly effective for LLMs, yet their application to vision-language models (VLMs) remains relatively underexplored. Existing VLM TTS met…