3 papers
cs.AI2026
Character Training for Risk-Averse Agents
Arav Dhoot, Punya Syon Pandey, Jamie Johnson +3
Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making dea…
cs.LG2026
Consistency Training Along the Transformer Stack
Sukrati Gautam, Neil Shah, Arav Dhoot +7
Consistency training encourages models to behave similarly across different contexts, and has shown promise for reducing misalignment. We broaden the scope of consistency training…
cs.LG2026
SemRep: Generative Code Representation Learning with Code Transformations
Weichen Li, Jiamin Song, Bogdan Alexandru Stoica +4
Code transformation is a foundational capability in the software development process, where its effectiveness relies on constructing a high-quality code representation to character…