5 papers
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
Kerem Zaman, Shashank Srivastava
Recent work, using the Biasing Features metric, labels a CoT as unfaithful if it omits a prompt-injected hint that affected the prediction. We argue this metric adopts a narrow not…
DiffVax: Optimization-Free Image Immunization Against Diffusion-Based Editing
Tarik Can Ozden, Ozgur Kara, Oguzhan Akcin +4
Current image immunization defense techniques against diffusion-based editing embed imperceptible noise into target images to disrupt editing models. However, these methods face sc…
A Causal Lens for Evaluating Faithfulness Metrics
Kerem Zaman, Shashank Srivastava
Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability. However, despite their plausibility, t…
INTERACT: Enabling Interactive, Question-Driven Learning in Large Language Models
Aum Kendapadi, Kerem Zaman, Rakesh R. Menon +1
Large language models (LLMs) excel at answering questions but remain passive learners-absorbing static data without the ability to question and refine knowledge. This paper explore…