14 papers
PrefDisco: Benchmarking Proactive Personalized Reasoning
Shuyue Stella Li, Avinandan Bose, Faeze Brahman +4
Current large language model (LLM) development treats task-solving and preference-alignment as separate challenges, optimizing first for objective correctness, then for alignment t…
Cold-Start Personalization via Training-Free Priors from Structured World Models
Avinandan Bose, Shuyue Stella Li, Faeze Brahman +6
Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available. The core challenge is a routing problem: each…
Sub-optimality of the Separation Principle for Quadratic Control from Bilinear Observations
Yahya Sattar, Sunmook Choi, Yassir Jedra +2
We consider the problem of controlling a linear dynamical system from bilinear observations with minimal quadratic cost. Despite the similarity of this problem to standard linear q…
Finite Sample Identification of Partially Observed Bilinear Dynamical Systems
Yahya Sattar, Yassir Jedra, Maryam Fazel +1
We consider the problem of learning a realization of a partially observed bilinear dynamical system (BLDS) from noisy input-output data. Given a single trajectory of input-output s…
Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics
Sunmook Choi, Yahya Sattar, Yassir Jedra +2
We study a nonstationary bandit problem where rewards depend on both actions and latent states, the latter governed by unknown linear dynamics. Crucially, the state dynamics also d…
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
Leo Boisvert, Mihir Bansal, Chandra Kiran Reddy Evuru +9
We present DoomArena, a security evaluation framework for AI agents. DoomArena is designed on three principles: 1) It is a plug-in framework and integrates easily into realistic ag…