4 papers · 1 filter
MEMO: Multimodal Evidence Memory Organization for Long-Horizon LLM Agents
Xian Gao, Jinpeng Wang, Jiacheng Ruan +3
Long-running LLM agents rely on external memory to store and reuse information beyond a single context window, yet there is a fundamental tension between the continuous accumulatio…
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
Jiacheng Ruan, Dan Jiang, Xian Gao +3
Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously ref…
MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
Xian Gao, Jiacheng Ruan, Zongyun Zhang +3
With the rapid growth of academic publications, peer review has become an essential yet time-consuming responsibility within the research community. Large Language Models (LLMs) ha…
ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews
Xian Gao, Jiacheng Ruan, Zongyun Zhang +3
Academic paper review is a critical yet time-consuming task within the research community. With the increasing volume of academic publications, automating the review process has be…