3 papers
cs.AI2026
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
Jie Wu, Ming Gong, Feixiang Cheng +1
Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite…
cs.CL2026
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
Shiping Yang, Jie Wu, Wenbiao Ding +7
Robustness has become a critical attribute for the deployment of RAG systems in real-world applications. Existing research focuses on robustness to explicit noise (e.g., document s…
cs.MA2026
Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems
Jie Wu, Ming Gong
We identify and formalize a novel security risk: Context-Fragmented Violations (CFVs) - a class of policy breaches where individual agent actions appear locally safe and reasonable…