Publications (40)
Superlatives in Context: Modeling the Implicit Semantics of Superlatives
Valentina Pyatkin, Bonnie Webber, Ido Dagan +1
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
Victoria Graf, Valentina Pyatkin, Nouha Dziri +2
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
Paul Röttger, Musashi Hinck, Valentin Hofmann +4
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
Linlu Qiu, Liwei Jiang, Ximing Lu +8
The Art of Saying No: Contextual Noncompliance in Language Models
Faeze Brahman, Sachin Kumar, Vidhisha Balachandran +11
Description-Based Text Similarity
Shauli Ravfogel, Valentina Pyatkin, Amir DN Cohen +2
Self-Directed Synthetic Dialogues and Revisions Technical Report
Nathan Lambert, Hailey Schoelkopf, Aaron Gokaslan +3
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
Jing-Jing Li, Joel Mire, Eve Fleisig +4
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
Lester James V. Miranda, Yizhong Wang, Yanai Elazar +6
Asking It All: Generating Contextualized Questions for any Semantic Role
Valentina Pyatkin, Paul Roit, Julian Michael +3
QASem Parsing: Text-to-text Modeling of QA-based Semantics
Ayal Klein, Eran Hirsch, Ron Eliav +3
Explicating the Implicit: Argument Detection Beyond Sentence Boundaries
Paul Roit, Aviv Slobodkin, Eran Hirsch +4
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
Jing-Jing Li, Valentina Pyatkin, Max Kleiman-Weiner +7
Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
Junyi Jessy Li, Yang Janet Liu, Kanishka Misra +2
Revisiting Sentence Union Generation as a Testbed for Text Consolidation
Eran Hirsch, Valentina Pyatkin, Ruben Wolhandler +3
The Possible, the Plausible, and the Desirable: Event-Based Modality Detection for Language Processing
Valentina Pyatkin, Shoval Sadde, Aynat Rubinstein +2
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Nathan Lambert, Jacob Morrison, Valentina Pyatkin +20
Promptly Predicting Structures: The Return of Inference
Maitrey Mehta, Valentina Pyatkin, Vivek Srikumar
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
Kavel Rao, Liwei Jiang, Valentina Pyatkin +5
Draw Me a Flower: Processing and Grounding Abstraction in Natural Language
Royi Lachmy, Valentina Pyatkin, Avshalom Manevich +1
Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
Hamish Ivison, Yizhong Wang, Valentina Pyatkin +8
ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral Situations
Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar +4
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
Taylor Sorensen, Liwei Jiang, Jena Hwang +10
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Hamish Ivison, Yizhong Wang, Jiacheng Liu +6
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
William Sheffield, Kanishka Misra, Valentina Pyatkin +3
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu +6
RewardBench: Evaluating Reward Models for Language Modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison +9
"You Are An Expert Linguistic Annotator": Limits of LLMs as Analyzers of Abstract Meaning Representation
Allyson Ettinger, Jena D. Hwang, Valentina Pyatkin +2
Olmo 3
Team Olmo, :, Allyson Ettinger +66
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
Alexis Limozin, Eduard Durech, Torsten Hoefler +2
Design Choices for Crowdsourcing Implicit Discourse Relations: Revealing the Biases Introduced by Task Design
Valentina Pyatkin, Frances Yung, Merel C. J. Scholman +3
Just-DREAM-about-it: Figurative Language Understanding with DREAM-FLUTE
Yuling Gu, Yao Fu, Valentina Pyatkin +3
Generalizing Verifiable Instruction Following
Valentina Pyatkin, Saumya Malik, Victoria Graf +5
2 OLMo 2 Furious
Team OLMo, Pete Walsh, Luca Soldaini +40
OLMo: Accelerating the Science of Language Models
Dirk Groeneveld, Iz Beltagy, Pete Walsh +40
Diverging Preferences: When do Annotators Disagree and do Models Know?
Michael JQ Zhang, Zhilin Wang, Jena D. Hwang +6
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
Paul Röttger, Valentin Hofmann, Valentina Pyatkin +4
QADiscourse -- Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines
Valentina Pyatkin, Ayal Klein, Reut Tsarfaty +1
RewardBench 2: Advancing Reward Model Evaluation
Saumya Malik, Valentina Pyatkin, Sander Land +4
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin +7