collaborators

17 papers

cs.CR2026

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

Pengfei He, Lesly Miculicich, Vishesh Sharma +5

Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt in…

cs.LG2026

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra +6

As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete de…

cs.CL2026

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

Gaotang Li, Bhavana Dalvi Mishra, Zifeng Wang +9

Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable…

cs.CV2026

ARD: Agentic Autoregressive Diffusion for Long Video Consistency

Do Xuan Long, Yale Song, Min-Yen Kan +2

Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse over long horizons. We present A…

cs.CL2026

PolicyBank: Evolving Policy Understanding for LLM Agents

Jihye Choi, Jinsung Yoon, Long T. Le +2

LLM agents operating under organizational policies must comply with authorization constraints typically specified in natural language. In practice, such specifications inevitably c…

cs.AI2026

TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems

Md Atik Ahamed, Mihir Parmar, Palash Goyal +7

We introduce TFRBench, the first benchmark designed to evaluate the reasoning capabilities of forecasting systems. Traditionally, time-series forecasting has been evaluated solely…