5 papers
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
Shutong Zhang, Dylan Zhou, Yinxiao Liu +3
The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. While large language models (LLM…
-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Haoran Zhang, Luxin Xu, Zhilin Wang +11
The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in…
FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains
Jiashuo Liu, Siyuan Chen, Zaiyuan Wang +38
Building upon FutureX, which established a live benchmark for general-purpose future prediction, this report introduces FutureX-Pro, including FutureX-Finance, FutureX-Retail, Futu…
LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
Liya Zhu, Peizhuang Cong, Jingzhe Ding +17
Large Language Models (LLMs) perform well on standard reasoning and question-answering benchmarks, yet such evaluations often fail to capture their ability to handle long-tail, exp…
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen +28
Future prediction is a complex task for LLM agents, requiring a high level of analytical thinking, information gathering, contextual understanding, and decision-making under uncert…