5 papers
Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection
Fatemeh Pesaran Zadeh, Seyeon Choi, Xing Han Lù +2
Large language models (LLMs) have enabled web agents that follow natural language goals through multi-step browser interactions. However, agents fine-tuned on specific trajectories…
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
Eunkyu Park, Wesley Hanwen Deng, Gunhee Kim +2
Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must perceive, understand, and judge all…
MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment
Eunkyu Park, Wesley Hanwen Deng, Cheyon Jin +8
Vision-Language Models (VLMs) continue to struggle to make morally salient judgments in multimodal and socially ambiguous contexts. Prior works typically rely on binary or pairwise…
Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan +4
Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We…
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
Eunkyu Park, Minyeong Kim, Gunhee Kim
Hallucinations pose a significant challenge to the reliability of large vision-language models, making their detection essential for ensuring accuracy in critical applications. Cur…