22 papers · 1 filter
A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models
Soham Dan, Himanshu Beniwal, Thomas Hartvigsen
Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural contexts. This survey synthesizes w…
Evaluating Temporal Consistency in Multi-Turn Language Models
Yash Kumar Atri, Steven L. Johnson, Tom Hartvigsen
Language models are increasingly deployed in interactive settings where users reason about facts over time rather than in isolation. In such scenarios, correct behavior requires mo…
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen +1
Speculative decoding has become the standard approach for accelerating Large Language Model (LLM) inference. It exploits a lossless draft-then-verify procedure to circumvent the la…
EDUMATH: Generating Standards-aligned Educational Math Word Problems
Bryan R. Christ, Penelope Molitz, Beau LeBlond +3
Math word problems (MWPs) are critical K-12 educational tools, and customizing them to students' interests and ability levels can enhance learning. However, teachers struggle to fi…
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
Youngwoo Kim, Himanshu Beniwal, Steven L. Johnson +1
Effective content moderation systems require explicit classification criteria, yet online communities like subreddits often operate with diverse, implicit standards. This work intr…
BEDTime: A Unified Benchmark for Automatically Describing Time Series
Medhasweta Sen, Zachary Gottesman, Jiaxing Qiu +3
Recent works propose complex multi-modal models that handle both time series and language, ultimately claiming high performance on complex tasks like time series reasoning and cros…