Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
Yang Zhou, Mingyu Zhao, Zhenting Wang +6
We present M^3-Bench, the first benchmark for evaluating multimodal tool use under the Model Context Protocol. The benchmark targets realistic, multi-hop and multi-threaded workflo…
cs.AI2025
APEER: Automatic Prompt Engineering Enhances Large Language Model Reranking
Can Jin, Hongwu Peng, Shiyu Zhao +7
Large Language Models (LLMs) have significantly enhanced Information Retrieval (IR) across various modules, such as reranking. Despite impressive performance, current zero-shot rel…