3 papers
cs.AI2026
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Lena Libon, Ben Rank, Jehyeok Yeon +5
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such…
cs.CV2026
Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems
Antonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii +7
Recent advances in image and video generation raise hopes that these models possess world modeling capabilities-the ability to generate realistic, physically plausible videos. This…
cs.CV2026
Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference
Agata Żywot, Iason Skylitsis, Thijmen Nijdam +4
Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without r…