activity
20242026
most citedMACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment

1 citations · 1 across the 2 of their papers we have counts for

collaborators

6 papers

cs.LG20261 cited

MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment

Ziyan Wang, Yali Du, Yudi Zhang +2

Offline Multi-agent Reinforcement Learning (MARL) is valuable in scenarios where online interaction is impractical or risky. While independent learning in MARL offers flexibility a…

cs.CL2026

Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving

Shunfeng Zheng, Yudi Zhang, Meng Fang +4

Retrieval-augmented generation (RAG) with foundation models has achieved strong performance across diverse tasks, but their capacity for expert-level reasoning-such as solving Olym…

cs.LG2025

Causality Meets Locality: Provably Generalizable and Scalable Policy Learning for Networked Systems

Hao Liang, Shuqing Shi, Yudi Zhang +2

Large-scale networked systems, such as traffic, power, and wireless grids, challenge reinforcement-learning agents with both scale and environment shifts. To address these challeng…

cs.AI2025

PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments

Olivier Schipper, Yudi Zhang, Yali Du +2

LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To…

cs.CL2025

Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?

Yudi Zhang, Lu Wang, Meng Fang +8

Distilling large language models (LLMs) typically involves transferring the teacher model's responses through supervised fine-tuning (SFT). However, this approach neglects the pote…

cs.AI2024

RuAG: Learned-rule-augmented Generation for Large Language Models

Yudi Zhang, Pei Xiao, Lu Wang +11

In-context learning (ICL) and Retrieval-Augmented Generation (RAG) have gained attention for their ability to enhance LLMs' reasoning by incorporating external knowledge but suffer…