activity
20152023
most citedDemonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP

53 citations · 186 across the 32 of their papers we have counts for

collaborators
Showing 2023 · cs.CLShow all

8 papers · 2 filters

cs.CL2023

ContextRef: Evaluating Referenceless Metrics For Image Description Generation

Elisa Kreiss, Eric Zelikman, Christopher Potts +1

Referenceless metrics (e.g., CLIPScore) use pretrained vision--language models to assess image descriptions directly without costly ground-truth reference texts. Such methods can f…

cs.CL2023

Rigorously Assessing Natural Language Explanations of Neurons

Jing Huang, Atticus Geiger, Karel D'Oosterlinck +2

Natural language is an appealing medium for explaining how large language models process and store information, but evaluating the faithfulness of such explanations is challenging.…

cs.CL2023

Context-VQA: Towards Context-Aware and Purposeful Visual Question Answering

Nandita Naik, Christopher Potts, Elisa Kreiss

Visual question answering (VQA) has the potential to make the Internet more accessible in an interactive way, allowing people who cannot see images to ask questions about them. How…

cs.CL2023★ 9 cited

ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning

Jingyuan Selena She, Christopher Potts, Samuel R. Bowman +1

A number of recent benchmarks seek to assess how well models handle natural language negation. However, these benchmarks lack the controlled example paradigms that would allow us t…

cs.CL2023★ 3 cited

MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions

Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning +2

The information stored in large language models (LLMs) falls out of date quickly, and retraining from scratch is often not an option. This has recently given rise to a range of tec…

cs.CL2023★ 8 cited

Interpretability at Scale: Identifying Causal Mechanisms in Alpaca

Zhengxuan Wu, Atticus Geiger, Thomas Icard +2

Obtaining human-interpretable explanations of large, general-purpose language models is an urgent goal for AI safety. However, it is just as important that our interpretability met…