Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences
Yiwen Jiang, Yang Deng, Stephanie Fong +9
Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from…
cs.AI2026
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…