Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows
Shivam Rawat, Lucie Flek
Agentic AI systems are increasingly being integrated into scientific workflows, yet their behavior under realistic conditions remains insufficiently understood. We evaluate CMBAgen…
cs.AI2025
Superalignment with Dynamic Human Values
Florian Mai, David Kaczér, Nicholas Kluge Corrêa +1
Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they…