8 citations · 14 across the 11 of their papers we have counts for
11 papers · 1 filter
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari +5
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuni…
Value Drifts: Tracing Value Alignment During LLM Post-Training
Mehar Bhatia, Shravan Nayak, Gaurav Kamath +4
As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to…
Infusing Theory of Mind into Socially Intelligent LLM Agents
EunJeong Hwang, Yuwei Yin, Giuseppe Carenini +2
Theory of Mind (ToM)-an understanding of the mental states of others-is a key aspect of human social intelligence, yet, chatbots and LLM-based social agents do not typically integr…
BottleHumor: Self-Informed Humor Explanation using the Information Bottleneck Principle
EunJeong Hwang, Peter West, Vered Shwartz
Humor is prevalent in online communications and it often relies on more than one modality (e.g., cartoons and memes). Interpreting humor in multimodal settings requires drawing on…
Bridging Information Gaps with Comprehensive Answers: Improving the Diversity and Informativeness of Follow-Up Questions
Zhe Liu, Taekyu Kang, Haoyu Wang +2
Generating diverse follow-up questions that uncover missing information remains challenging for conversational agents, particularly when they run on small, locally hosted models. T…
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin +8
Robust, diverse, and challenging cultural knowledge benchmarks are essential for measuring our progress towards making LMs that are helpful across diverse cultures. We introduce Cu…