2 papers
cs.AI2026
AgentCPM-Explore: Realizing Long-Horizon Deep Exploration for Edge-Scale Agents
Haotian Chen, Xin Cong, Shengda Fan +16
While Large Language Model (LLM)-based agents have shown remarkable potential for solving complex tasks, existing systems remain heavily reliant on large-scale models, leaving the…
cs.CL2025
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
Xuanwen Ding, Chengjun Pan, Zejun Li +3
Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. T…