4 papers
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
Xiang Feng, Jiawei Zhou, Zhangfeng Huang +6
Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end-to-end answer ac…
Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation
Jiawei Zhou, Chi Zhang, Xiang Feng +6
We present Omni-I2C, a comprehensive benchmark designed to evaluate the capability of Large Multimodal Models (LMMs) in converting complex, structured digital graphics into executa…
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
Xiang Feng, Wentao Jiang, Zengmao Wang +5
The application of large language models (LLMs) in the medical field has garnered significant attention, yet their reasoning capabilities in more specialized domains like anesthesi…
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
Wentao Jiang, Xiang Feng, Zengmao Wang +5
Reinforcement learning (RL) is emerging as a powerful paradigm for enabling large language models (LLMs) to perform complex reasoning tasks. Recent advances indicate that integrati…