2 papers
cs.LG2025
Average-DICE: Stationary Distribution Correction by Regression
Fengdi Che, Bryan Chan, Chen Ma +1
Off-policy policy evaluation (OPE), an essential component of reinforcement learning, has long suffered from stationary state distribution mismatch, undermining both stability and…
cs.LG2025
Toward Understanding In-context vs. In-weight Learning
Bryan Chan, Xinyi Chen, András György +1
It has recently been demonstrated empirically that in-context learning emerges in transformers when certain distributional properties are present in the training data, but this abi…