5 papers
Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs
Jinbo Liu, Defu Cao, Yifei Wei +6
Graph topology is a fundamental determinant of memory leakage in multi-agent LLM systems, yet its effects remain poorly quantified. We introduce MAMA (Multi-Agent Memory Attack), a…
Counterfactual Trace Auditing of LLM Agent Skills
Xiaolin Zhou, Jinbo Liu, Li Li +2
Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchmarks report only pass rate befor…
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Xiaolin Zhou, Aojie Yuan, Zheng Luo +12
Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user…
Conversational Time Series Foundation Models: Towards Explainable and Effective Forecasting
Defu Cao, Michael Gee, Jinbo Liu +4
The proliferation of time series foundation models has created a landscape where no single method achieves consistent superiority, framing the central challenge not as finding the…
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
Wen Ye, Jinbo Liu, Defu Cao +2
The rapid advancement of Large Language Models (LLMs) has sparked growing interest in their application to time series analysis tasks. However, their ability to perform complex rea…