1 citations · 1 across the 11 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
World-Time Compute with Verified Code World Models
James Schwoebel, Ingrida Semenec, Jenia Rousseva +8
LLMs generalize across a domain only after seeing many real, labeled examples, which most domains lack. We study a way to manufacture it cheaply. When a domain's dynamics can be wr…
cs.LG2026
Manifold of Failure: Behavioral Attraction Basins in Language Models
Sarthak Munshi, Manish Bhatt, Vineeth Sai Narajala +4
While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehensive understanding of AI safety r…