Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
Xinyu Lu, Tianshu Wang, Pengbo Wang +8
Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next-level capability: whether m…
cs.AI2026
Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games
Mingyuan Fan, Weiguang Han, Daixin Wang +3
We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Wherein, LLMs receive only the ta…
cs.AI2025
Making Large Language Models Better Knowledge Miners for Online Marketing with Progressive Prompting Augmentation
Chunjing Gan, Dan Yang, Binbin Hu +6
Nowadays, the rapid development of mobile economy has promoted the flourishing of online marketing campaigns, whose success greatly hinges on the efficient matching between user pr…