19 papers
Vector Symbolic Policy Gradient
Ryozo Masukawa, Sanggeon Yun, SungHeon Jeong +6
We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to t…
Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)
Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, h…
Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes
Zuyuan Zhang, Yongshan Chen, Mahdi Imani +1
The paper defines the smallest memory representation needed for a class of partially observable decision processes called holonomy-cover decision processes, builds a stable quotien…
The Topology of Ill-Posed Questions: Persistent Homology for Detection and Steering in LLMs
Guangyu Jiang, Sizhe Tang, Mahdi Imani +1
Ill-posed questions, including ambiguous, underspecified, or contradictory queries, may admit no valid answer or multiple plausible answers, posing a challenge for large language m…
FedQHD: Closed-Form Function-Space Federated Reinforcement Learning
Yuchen Hou, Yongshan Chen, Zhuowen Zou +4
Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories. However, FedAvg-style para…
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
Zuyuan Zhang, Sizhe Tang, Mahdi Imani +1
General-sum multi-agent learning is often governed by a stacked update field in which each agent's policy update changes the optimization landscape faced by the others. This coupli…