21 papers
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
Yiyao Zhang, Diksha Goel, Hussain Ahmad +2
An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accur…
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
Yiyao Zhang, Diksha Goel, Hussain Ahmad +2
A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both halves of that requirement wit…
DEFENGRAPH: Knowledge Graph-Enhanced LLMs for Blue Team Cyber Defense
Zhen Wang, Kristen Moore, Qin Wang +7
Large Language Models (LLMs) show promise for supporting decision-making in cybersecurity, but their reliability in high-stakes, time-evolving environments remains limited due to h…
Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents
Diksha Goel, Kristen Moore, Jeff Wang +2
Reinforcement Learning (RL) agents are increasingly used to simulate sophisticated cyberattacks, but their decision-making processes remain opaque, hindering trust, debugging, and…
AgenticVM: Agentic AI for Adaptive Software Vulnerability Management
Asrul Arifin, Hussain Ahmad, Yiyao Zhang +1
As software systems grow in scale and complexity, vulnerability management is increasingly strained by high alert volumes, fragmented toolchains, and manual triage processes. We in…
Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning
Yiyao Zhang, Diksha Goel, Hussain Ahmad
Autonomous agents are increasingly deployed in both offensive and defensive cyber operations, creating high-speed, closed-loop interactions in critical infrastructure environments.…