6 citations · 14 across the 34 of their papers we have counts for
5 papers · 1 filter
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Sinie van der Ben, Raphaël Baur, Yannick Metz +1
Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirr…
Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain
Corentin Royer, Debarun Bhattacharjya, Gaetano Rossiello +2
Multi-step reasoning improves the capabilities of large language models (LLMs) but increases the risk of errors propagating through intermediate steps. Process reward models (PRMs)…
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
Chenfei Xiong, Jingwei Ni, Yu Fan +10
We introduce Co-DETECT (Collaborative Discovery of Edge cases in TExt ClassificaTion), a novel mixed-initiative annotation framework that integrates human expertise with automatic…
Concept-Level Explainability for Auditing & Steering LLM Responses
Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady
As large language models (LLMs) become widely deployed, concerns about their safety and alignment grow. An approach to steer LLM behavior, such as mitigating biases or defending ag…
LayerFlow: Layer-wise Exploration of LLM Embeddings using Uncertainty-aware Interlinked Projections
Rita Sevastjanova, Robin Gerling, Thilo Spinner +1
Large language models (LLMs) represent words through contextual word embeddings encoding different language properties like semantics and syntax. Understanding these properties is…