4 papers
EARS: Explanatory Abstention for Reliable Sub-Agent Modeling in Large-scale Multi-Agent Systems
Shuang Xie, Yunan Lu, Han Li +1
In large-scale enterprise settings, centralized multi-agent systems (MAS) are increasingly adopted, in which a coordinator delegates user requests to lightweight, domain-specialize…
Mitigating Hallucinations in Large Language Models Via Decoder Layer Skipping
Hanze Li, Jinhao You, Yichen Guo +3
Large Language Models (LLMs) have achieved strong performance across diverse natural language tasks, yet their outputs often suffer from hallucinations -- content that is misaligne…
SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents
Han Li, Vibhor Malik, Zahra Zanjani Foumani +17
A/B testing remains the gold standard for evaluating modifications to e-commerce storefronts, yet it diverts traffic, requires weeks to reach statistical significance, and risks de…
ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents
Chinmay Savadikar, Mingyu Zhao, Yuanzheng Zhu +5
Developing and evaluating e-commerce web agents requires environments that preserve meaningful task structure while enabling controllable, reproducible, and scalable scientific com…