collaborators

34 papers

cs.AI2026

Aligning Language Model Benchmarks with Pairwise Preferences

Marco Gutierrez, Xinyi Leng, Hannah Cyberey +3

Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real…

cs.CL2026

A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

Soham Dan, Himanshu Beniwal, Thomas Hartvigsen

Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural contexts. This survey synthesizes w…

cs.AI2026

Inferring Events from Time Series using Language Models

Mingtian Tan, Mike A. Merrill, Zack Gottesman +3

A common goal in analyzing time series data is to understand how events cause observed variations. We study whether Large Language Models (LLMs) can infer natural language events a…

cs.LG2026

LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws

Xu Ouyang, Deyi Liu, Yuhang Cai +5

Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and q…

cs.CL2026

Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?

Natalie Seah, Danielle S. Bitterman, Daphna Spiegel +1

Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where clinicians must clearly and…

cs.CV2026

Test-Time Hinting for Black-Box Vision-Language Models

Kaihua Hou, Abhijith Varma Mudunuri, Jiaxing Qiu +3

Test-time scaling (TTS) methods have proven highly effective for LLMs, yet their application to vision-language models (VLMs) remains relatively underexplored. Existing VLM TTS met…