16 papers
ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation
Joseph Marvin Imperial, Junhong Liang, Belal Shoer +9
When a text is translated, does the translation retain the complexity of the original? We introduce ComplexityMT, a new challenge for assessing how text complexity and machine tran…
Linguistic Productivity in Large Language Models: Models Coerce, but do not Preempt
Claire Bonial, Claire Benet Post, Laura Michaelis +1
Usage-based theories of grammars posit that creative productivity of the structures of language is both bolstered and constrained by two distinct frequency signals: entrenchment, s…
Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions
Wesley Scivetti, Melissa Torgbi, Austin Blodgett +4
The web-scale of pretraining data has created an important evaluation challenge: to disentangle linguistic competence on cases well-represented in pretraining data from generalizat…
Safer Policy Compliance with Dynamic Epistemic Fallback
Joseph Marvin Imperial, Harish Tayyar Madabushi
Humans develop a series of cognitive defenses, known as epistemic vigilance, to combat risks of deception and misinformation from everyday interactions. Developing safeguards for L…
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
Jingcheng Niu, Subhabrata Dutta, Ahmed Elshabrawy +2
Large-scale Transformer language models (LMs) trained solely on next-token prediction with web-scale data can solve a wide range of tasks after seeing just a few examples. The mech…
Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
Joseph Marvin Imperial, Harish Tayyar Madabushi
Policy compliance assessment is a fundamental task of evaluating whether an input case strictly complies with a set of human-defined rules, more generally known as policies. In pra…