2 papers
cs.AI2026
CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges
Zi-Han Wang, Lam Nguyen, Zhengyang Zhao +4
The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success o…
cs.HC2026
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
Yuanbo Tang, Huaze Tang, Tingyu Cao +6
Proactive agents that anticipate user intentions without explicit prompts represent a significant evolution in human-AI interaction, promising to reduce cognitive load and streamli…