1 paper · 1 filter
Darrin O' Brien, Dhikshith Gajulapalli, Eric Xia
Results in interpretability suggest that large vision and language models learn implicit linear encodings when models are biased by in-context prompting. However, the existence of…