4 papers · 1 filter
Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
Sungjun Lim, Heedong Kim, Andrew Lee +1
Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and…
Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach
Youngjun Choi, Joonseong Kang, Sungjun Lim +1
Data valuation has become central in the era of data-centric AI. It drives efficient training pipelines and enables objective pricing in data markets by assigning a numeric value t…
Semi-Supervised Preference Optimization with Limited Feedback
Seonggyun Lee, Sungjun Lim, Seojin Park +2
The field of preference optimization has made outstanding contributions to the alignment of language models with human preferences. Despite these advancements, recent methods still…
Uncertainty-driven Embedding Convolution
Sungjun Lim, Kangjun Noh, Youngjun Choi +2
Text embeddings are essential components in modern NLP pipelines. Although numerous embedding models have been proposed, no single model consistently dominates across domains and t…