activity
20192025
most citedDynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking

16 citations · 28 across the 7 of their papers we have counts for

collaborators

12 papers

cs.AI2025

Humanline: Online Alignment as Perceptual Loss

Sijia Liu, Niklas Muennighoff, Kawin Ethayarajh

Online alignment (e.g., GRPO) is generally more performant than offline alignment (e.g., DPO) -- but why? Drawing on prospect theory from behavioral economics, we propose a human-c…

cs.CL2022

Richer Countries and Richer Representations

Kaitlyn Zhou, Kawin Ethayarajh, Dan Jurafsky

We examine whether some countries are more richly represented in embedding space than others. We find that countries whose names occur with low frequency in training corpora are mo…

cs.CL20221 cited

Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words

Kaitlyn Zhou, Kawin Ethayarajh, Dallas Card +1

Cosine similarity of contextual embeddings is used in many NLP tasks (e.g., QA, IR, MT) and metrics (e.g., BERTScore). Here, we uncover systematic ways in which word similarities e…

cs.CL2021

Conditional probing: measuring usable information beyond a baseline

John Hewitt, Kawin Ethayarajh, Percy Liang +1

Probing experiments investigate the extent to which neural representations make properties -- like part-of-speech -- predictable. One suggests that a representation encodes a prope…

cs.CL20213 cited

Attention Flows are Shapley Value Explanations

Kawin Ethayarajh, Dan Jurafsky

Shapley Values, a solution to the credit assignment problem in cooperative game theory, are a popular type of explanation in machine learning, having been used to explain the impor…

cs.CL202116 cited

Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking

Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush +6

We introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform e…