Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
World Models in Pieces: Structural Certification for General Agents
Yikai Lu, Yifei Wu, Xinyu Lu +1
In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees…
cs.AI2026
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent
Zeyao Du, Tong Li, Yanci Zhang +1
As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or re…
cs.AI2025
Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling
Fanzeng Xia, Yidong Luo, Tinko Sebastian Bartels +2
Recent research has highlighted that Large Language Models (LLMs), even when trained to generate extended long reasoning steps, still face significant challenges on hard reasoning…