11 papers
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
Hao Yan, Yuliang Liu, Xingchen Liu +5
Existing Multimodal Large Language Models (MLLMs) suffer from significant performance degradation on the long document understanding task as document length increases. This stems f…
Benchmarking Real-Time Question Answering via Executable Code Workflows
Wenjie Zhou, Yuan Gao, Xin Zhou +5
Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and ther…
Construction of Knowledge Graph based on Language Model
Qiubai Zhu, Qingwang Wang, Haibin Yuan +2
Knowledge Graph (KG) can effectively integrate valuable information from massive data, and thus has been rapidly developed and widely used in many fields. Traditional KG constructi…
STRIDE: Strategic Iterative Decision-Making for Retrieval-Augmented Multi-Hop Question Answering
Wei Chen, Lili Zhao, Zhi Zheng +2
Multi-hop question answering (MHQA) enables accurate answers to complex queries by retrieving and reasoning over evidence dispersed across multiple documents. Existing MHQA approac…
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
Sirry Chen, Jieyi Wang, Wei Chen +1
Medical consultations are intrinsically speech-centric. However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly. Recent advance…
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
Haoyi Sun, Xiaoxiao Wang, Ning Mao +5
Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challenges for deployment in resour…