2 papers
cs.CL2026
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Xiaohongshu Inc, Xiaohongshu Dots Studio, Evolvent AI
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments.…
cs.CL2026
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
Xiaohongshu Inc
LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to…