5 papers
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
Yuming Yang, Xiao Sun, Yuanwei Zou +6
Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simu…
Lung-R1: A Knowledge Graph-Guided LLM for Pulmonary Diagnostic Reasoning
Haoyang Zeng, Yuanxi Fu, Rongzhen Li +11
Diagnosing pulmonary diseases requires integrating heterogeneous evidence amid phenotypic variability and cross-disease overlap. Although large language models (LLMs) have shown pr…
Do Models See in Line with Human Vision? Probing the Correspondence Between LVLM Representations and EEG Signals
Xin Xiao, Yang Lei, Haoyang Zeng +6
Large Vision Language Models (LVLMs) exhibit strong visual understanding and reasoning abilities. However, whether their internal representations reflect human visual cognition is…
ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios
Changzai Pan, Jie Zhang, Kaiwen Wei +15
Recent advancements in Large Language Models (LLMs) have significantly catalyzed table-based question answering (TableQA). However, existing TableQA benchmarks often overlook the i…
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
Kaiwen Wei, Xiao Liu, Jie Zhang +11
Multimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence, and numerous video-based…