14 papers
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
Jiamin Chen, Yidi Wu, Qiexiang Wang +6
Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than construct…
DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation
Jiamin Chen, Qianben Chen, Jiawen Zhang +5
Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinematic control, audio, and cros…
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
Qian Chen, Xianyin Zhang, Yanzhi Liu +3
The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual an…
GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
Hao-Xiang Xu, Chong Deng, Jiaqing Liu +5
Large Language Models (LLMs) extend their capabilities through function-calling (FC), which relies on training data with high quality, diversity, and broad coverage of scenario. Ho…
Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions
Yue Xu, Qian Chen, Zizhan Ma +5
Large language models have enabled agentic systems that reason, plan, and interact with tools and environments to accomplish complex tasks. As these agents operate over extended in…
Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
Qianben Chen, Tianrui Qin, King Zhu +21
Recent deep research agents primarily improve performance by scaling reasoning depth, but this leads to high inference cost and latency in search-intensive scenarios. Moreover, gen…