3 papers
cs.AI2025
Recontextualization Mitigates Specification Gaming without Modifying the Specification
Ariana Azarbal, Victor Gillioz, Vladimir Ivanov +6
Developers often struggle to specify correct training labels and rewards. Perhaps they don't need to. We propose recontextualization, which reduces how often language models "game"…
cs.LG2025
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
Nevan Wichers, Aram Ebtekar, Ariana Azarbal +8
Large language models are sometimes trained with imperfect oversight signals, leading to undesired behaviors such as reward hacking and sycophancy. Improving oversight quality can…
cs.CL2025
Dynamic Relation Inference via Verb Embeddings
Omri Suissa, Muhiim Ali, Ariana Azarbal +2
CLIP has demonstrated exceptional image-text matching capabilities due to its training on contrastive learning tasks. Past research has suggested that whereas CLIP effectively matc…