works on

From the 1 of 9 linked papers with an AI index.

most citedMACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment

1 citations · 1 across the 5 of their papers we have counts for

collaborators

9 papers

cs.MA2026

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration

Xinxing Ren, Qianbo Zang, Ziyan Wang +4

The paper introduces AgentRadio, an asynchronous message‑passing layer that enables multiple LLM coding agents to stay passively aware of each other's findings during long‑horizon…

cs.RO2026

Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

Ziyan Wang, Tan Xiang, Peng Chen +1

A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observable closed-loop environments. I…

cs.LG20261 cited

MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment

Ziyan Wang, Yali Du, Yudi Zhang +2

Offline Multi-agent Reinforcement Learning (MARL) is valuable in scenarios where online interaction is impractical or risky. While independent learning in MARL offers flexibility a…

cs.LG2026

Fisher Decorator: Refining Flow Policy via a Local Transport Map

Xiaoyuan Cheng, Haoyu Wang, Wenxuan Yuan +4

Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However, they still face critical t…

cs.AI2026

MEMENTO: Teaching LLMs to Manage Their Own Context

Vasilis Kontonis, Yuchen Zeng, Shivam Garg +7

Reasoning models think in long, unstructured streams with no mechanism for compressing or organizing their own intermediate state. We introduce MEMENTO: a method that teaches model…

cs.LG2025

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

Zhicheng Zhang, Ziyan Wang, Yali Du +1

Developing effective instruction-following policies in reinforcement learning remains challenging due to the reliance on extensive human-labeled instruction datasets and the diffic…