1 citations · 1 across the 4 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
Huaixiao Tou, Ying Zeng, Yuemeng Li +6
We present ShoppingComp, a challenging real-world benchmark for comprehensively evaluating LLM-powered shopping agents on three core capabilities: precise product retrieval, expert…
cs.CL2025★ 1 cited
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Minghao Li, Ying Zeng, Zhihao Cheng +2
The advent of Deep Research agents has substantially reduced the time required for conducting extensive research tasks. However, these tasks inherently demand rigorous standards of…