chain-of-thought 1evaluation benchmarks 1large language models 1model reliability 1multimodal models 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Zeyu Chen, Huanjin Yao, Ziwang Zhao +1
The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…
cs.CL2026
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
Zeyuan Chen, Ziqing Yang, Yihan Ma +2
As academic submissions grow, the traditional peer review process struggles to keep up, raising concerns about quality and fairness. A trend of using large language models (LLMs) f…
cs.CR2026
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
Yage Zhang, Yukun Jiang, Zeyuan Chen +3
Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions. These limitations dri…