6 papers
START: Spatial and Textual Learning for Chart Understanding
Zhuoming Liu, Xiaofeng Gao, Feiyang Niu +3
Chart understanding is crucial for deploying multimodal large language models (MLLMs) in real-world scenarios such as analyzing scientific papers and technical reports. Unlike natu…
SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning
Xuchen Li, Ruitao Wu, Xuanbo Liu +17
Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcr…
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
Xukai Wang, Xuanbo Liu, Mingrui Chen +16
With the advancement of powerful large-scale reasoning models, effectively evaluating the reasoning capabilities of these models has become increasingly important. However, existin…
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
Xikai Zhang, Bo Wang, Likang Xiao +4
Although large language models (LLMs) have made significant strides across various tasks, they still face significant challenges in complex reasoning and planning. For example, eve…
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
Xiangfeng Wang, Xiao Li, Yadong Wei +8
The rapid growth of online video content, especially on short video platforms, has created a growing demand for efficient video editing techniques that can condense long-form video…
Research on Graph-Retrieval Augmented Generation Based on Historical Text Knowledge Graphs
Yang Fan, Zhang Qi, Xing Wenqian +2
This article addresses domain knowledge gaps in general large language models for historical text analysis in the context of computational humanities and AIGC technology. We propos…