4 papers
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
Jun Feng, Zixin Wang, Zhentao Zhang +5
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in visual mathematical reasoning across various existing benchmarks. However, these benchmarks ar…
DeepShop: A Benchmark for Deep Research Shopping Agents
Yougang Lyu, Xiaoyu Zhang, Lingyong Yan +3
Web agents for online shopping have shown great promise in automating user interactions across e-commerce platforms. Benchmarks for assessing such agents do not reflect the complex…
Unifying Search and Recommendation with Dual-View Representation Learning in a Generative Paradigm
Jujia Zhao, Wenjie Wang, Chen Xu +3
Recommender systems and search engines serve as foundational elements of online platforms, with the former delivering information proactively and the latter enabling users to seek…
Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs
Zixiao Wang, Duzhen Zhang, Ishita Agrawal +3
Previous approaches to persona simulation large language models (LLMs) have typically relied on learning basic biographical information, or using limited role-play dialogue dataset…