Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce
Shimaa Ahmed, Yiwei Cai, Mohsen Minaei +1
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchant…
cs.AI2026
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
Liang Wang, Junpeng Wang, Chin-chia Michael Yeh +6
Large Language Models (LLMs) are increasingly used as evaluators of reasoning quality, yet their reliability and bias in payments-risk settings remain poorly understood. We introdu…