5 papers
Understanding Impact of Human Feedback via Influence Functions
Taywon Min, Haeone Lee, Yongchan Kwon +1
In Reinforcement Learning from Human Feedback (RLHF), it is crucial to learn suitable reward models from human feedback to align large language models (LLMs) with human intentions.…
Newfluence: Boosting Model interpretability and Understanding in High Dimensions
Haolin Zou, Arnab Auddy, Yongchan Kwon +2
The increasing complexity of machine learning (ML) and artificial intelligence (AI) models has created a pressing need for tools that help scientists, engineers, and policymakers i…
Certified Data Removal Under High-dimensional Settings
Haolin Zou, Arnab Auddy, Yongchan Kwon +2
Machine unlearning focuses on the computationally efficient removal of specific training data from trained models, ensuring that the influence of forgotten data is effectively elim…
Distributionally Robust Instrumental Variables Estimation
Zhaonan Qu, Yongchan Kwon
Instrumental variables (IV) estimation is a fundamental method in econometrics and statistics for estimating causal effects in the presence of unobserved confounding. However, chal…
Group Shapley Value and Counterfactual Simulations in a Structural Model
Yongchan Kwon, Sokbae Lee, Guillaume A. Pouliot
We propose a variant of the Shapley value, the group Shapley value, to interpret counterfactual simulations in structural economic models by quantifying the importance of different…