3 papers
cs.LG2025
On Transportability for Structural Causal Bandits
Min Woo Park, Sanghack Lee
Intelligent agents equipped with causal knowledge can optimize their action spaces to avoid unnecessary exploration. The structural causal bandit framework provides a graphical cha…
cs.CL2025
Mitigating Length Bias in RLHF through a Causal Lens
Hyeonji Kim, Sujeong Oh, Sanghack Lee
Reinforcement learning from human feedback (RLHF) is widely used to align large language models (LLMs) with human preferences. However, RLHF-trained reward models often exhibit len…
cs.LG2025
On Predicting Post-Click Conversion Rate via Counterfactual Inference
Junhyung Ahn, Sanghack Lee
Accurately predicting conversion rate (CVR) is essential in various recommendation domains such as online advertising systems and e-commerce. These systems utilize user interaction…