2 citations · 4 across the 4 of their papers we have counts for
4 papers
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
Darshan Deshpande, Selvan Sunitha Ravi, Sky CH-Wang +3
The LLM-as-judge paradigm is increasingly being adopted for automated evaluation of model outputs. While LLM judges have shown promise on constrained evaluation tasks, closed sourc…
Whose ChatGPT? Unveiling Real-World Educational Inequalities Introduced by Large Language Models
Renzhe Yu, Zhen Xu, Sky CH-Wang +1
The universal availability of ChatGPT and other similar tools since late 2022 has prompted tremendous public excitement and experimental effort about the potential of large languag…
NormDial: A Comparable Bilingual Synthetic Dialog Dataset for Modeling Social Norm Adherence and Violation
Oliver Li, Mallika Subramanian, Arkadiy Saakyan +2
Social norms fundamentally shape interpersonal communication. We present NormDial, a high-quality dyadic dialogue dataset with turn-by-turn annotations of social norm adherences an…
Toward a Critical Toponymy Framework for Named Entity Recognition: A Case Study of Airbnb in New York City
Mikael Brunila, Jack LaViolette, Sky CH-Wang +3
Critical toponymy examines the dynamics of power, capital, and resistance through place names and the sites to which they refer. Studies here have traditionally focused on the sema…