8 papers
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
Huanyao Zhang, Jiepeng Zhou, Runhao Zhao +12
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive…
GenOM: Ontology Matching with Description Generation and Large Language Model
Yiping Song, Jiaoyan Chen, Renate A. Schmidt
Ontology matching (OM) plays an essential role in enabling semantic interoperability and integration across heterogeneous knowledge sources, particularly in the biomedical domain w…
BrowseComp-: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents
Huanyao Zhang, Jiepeng Zhou, Bo Li +22
Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimod…
Aggregation Queries over Unstructured Text: Benchmark and Agentic Method
Haojia Zhu, Qinyuan Xu, Haoyu Li +4
Aggregation query over free text is a long-standing yet underexplored problem. Unlike ordinary question answering, aggregate queries require exhaustive evidence collection and syst…
Ontology-Enhanced Knowledge Graph Completion using Large Language Models
Wenbin Guo, Xin Wang, Jiaoyan Chen +2
Large Language Models (LLMs) have been extensively adopted in Knowledge Graph Completion (KGC), showcasing significant research advancements. However, as black-box models driven by…
L3A: Label-Augmented Analytic Adaptation for Multi-Label Class Incremental Learning
Xiang Zhang, Run He, Jiao Chen +5
Class-incremental learning (CIL) enables models to learn new classes continually without forgetting previously acquired knowledge. Multi-label CIL (MLCIL) extends CIL to a real-wor…