37 citations · 69 across the 11 of their papers we have counts for
11 papers · 1 filter
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
Yu Ying Chiu, Liwei Jiang, Maria Antoniak +7
Frontier large language models (LLMs) are developed by researchers and practitioners with skewed cultural backgrounds and on datasets with skewed sources. However, LLMs' (lack of)…
CASE: Commonsense-Augmented Score with an Expanded Answer Space
Wenkai Chen, Sahithya Ravi, Vered Shwartz
LLMs have demonstrated impressive zero-shot performance on NLP tasks thanks to the knowledge they acquired in their training. In multiple-choice QA tasks, the LM probabilities are…
Automatic Evaluation of Generative Models with Instruction Tuning
Shuhaib Mehri, Vered Shwartz
Automatic evaluation of natural language generation has long been an elusive goal in NLP.A recent paradigm fine-tunes pre-trained language models to emulate human judgements for a…
GD-COMET: A Geo-Diverse Commonsense Inference Model
Mehar Bhatia, Vered Shwartz
With the increasing integration of AI into everyday life, it's becoming crucial to design AI systems that serve users from diverse backgrounds by making them culturally aware. In t…
Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi +5
The escalating debate on AI's capabilities warrants developing reliable metrics to assess machine "intelligence". Recently, many anecdotal examples were used to suggest that newer…
From chocolate bunny to chocolate crocodile: Do Language Models Understand Noun Compounds?
Jordan Coil, Vered Shwartz
Noun compound interpretation is the task of expressing a noun compound (e.g. chocolate bunny) in a free-text paraphrase that makes the relationship between the constituent nouns ex…