most citedBeware the Rationalization Trap! When Language Model Explainability Diverges from our Mental Models of Language

5 citations · 15 across the 7 of their papers we have counts for

collaborators

20 papers

cs.HC2025

Dia-Lingle: A Gamified Interface for Dialectal Data Collection

Jiugeng Sun, Rita Sevastjanova, Sina Ahmadi +2

Dialects suffer from the scarcity of computational textual resources as they exist predominantly in spoken rather than written form and exhibit remarkable geographical diversity. C…

cs.HC20256 cited

Beyond Quantification: Navigating Uncertainty in Professional AI Systems

Sylvie Delacroix, Diana Robinson, Umang Bhatt +12

The growing integration of large language models across professional domains transforms how experts make critical decisions in healthcare, education, and law. While significant res…

cs.HC20251 cited

DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition

Danqing Shi, Furui Cheng, Tino Weinkauf +2

Human preferences are widely used to align large language models (LLMs) through methods such as reinforcement learning from human feedback (RLHF). However, the current user interfa…

cs.CG2025

Explainable Mapper: Charting LLM Embedding Spaces Using Perturbation-Based Explanation and Verification Agents

Xinyuan Yan, Rita Sevastjanova, Sinie van der Ben +2

Large language models (LLMs) produce high-dimensional embeddings that capture rich semantic and syntactic relationships between words, sentences, and concepts. Investigating the to…

cs.CL20251 cited

Concept-Level Explainability for Auditing & Steering LLM Responses

Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady

As large language models (LLMs) become widely deployed, concerns about their safety and alignment grow. An approach to steer LLM behavior, such as mitigating biases or defending ag…

cs.CL2025

LayerFlow: Layer-wise Exploration of LLM Embeddings using Uncertainty-aware Interlinked Projections

Rita Sevastjanova, Robin Gerling, Thilo Spinner +1

Large language models (LLMs) represent words through contextual word embeddings encoding different language properties like semantics and syntax. Understanding these properties is…