52 citations · 66 across the 15 of their papers we have counts for
16 papers
XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?
Akhila Yerukola, Jena D. Hwang, Mingqian Zheng +5
When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g., "How do I parse XML with regex?"). In such cases, often referred to as the XY-prob…
NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
Akhila Yerukola, Fabrice Y Harel-Canada, Simran Khanuja +5
AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior work on cultural understanding evaluates AI systems on text-only settin…
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
Jocelyn Shen, Akhila Yerukola, Xuhui Zhou +3
Conversational breakdowns in close relationships are deeply shaped by personal histories and emotional context, yet most NLP research treats conflict detection as a general task, o…
Out of Style: RAG's Fragility to Linguistic Variation
Tianyu Cao, Neel Bhandari, Akhila Yerukola +2
Despite the impressive performance of Retrieval-augmented Generation (RAG) systems across various NLP benchmarks, their robustness in handling real-world user-LLM interaction queri…
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
Priyanshu Kumar, Devansh Jain, Akhila Yerukola +4
Truly multilingual safety moderation efforts for Large Language Models (LLMs) have been hindered by a narrow focus on a small set of languages (e.g., English, Chinese) as well as a…
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
Akhila Yerukola, Saadia Gabriel, Nanyun Peng +1
Gestures are an integral part of non-verbal communication, with meanings that vary across cultures, and misinterpretations that can have serious social and diplomatic consequences.…