6 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 1 cited
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
Karel D'Oosterlinck, Winnie Xu, Chris Develder +5
Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes…
cs.CL2022★ 6 cited
Causal Proxy Models for Concept-Based Model Explanations
Zhengxuan Wu, Karel D'Oosterlinck, Atticus Geiger +2
Explainability methods for NLP systems encounter a version of the fundamental problem of causal inference: for a given ground-truth input text, we never truly observe the counterfa…