5 papers
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
Jie Wu, Ming Gong, Feixiang Cheng +1
Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite…
Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems
Jie Wu, Ming Gong
We identify and formalize a novel security risk: Context-Fragmented Violations (CFVs) - a class of policy breaches where individual agent actions appear locally safe and reasonable…
Policy-Invisible Violations in LLM-Based Agents
Jie Wu, Ming Gong
LLM-based agents can execute actions that are syntactically valid, user-sanctioned, and semantically appropriate, yet still violate organizational policy because the facts needed f…
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
Haoxin Wang, Xianhan Peng, Xucheng Huang +5
In this paper, we introduce ECom-Bench, the first benchmark framework for evaluating LLM agent with multimodal capabilities in the e-commerce customer support domain. ECom-Bench fe…
MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
Ming Gong, Xucheng Huang, Chenghan Yang +4
Recent advances in large language models (LLMs) have enabled new applications in e-commerce customer service. However, their capabilities remain constrained in complex, multimodal…