2 papers
cs.AI2026
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
Zechun Niu, Yukun Zhao, Jiaxin Zhang +12
Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or…
cs.CL2026
SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay
Guilin Li, Jiaxing Zhang, Matthias Hwai Yong Tan +2
Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful acti…