7 papers
LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment
Fangyu Hao, Jiayu Yang, Yifan Zhu +14
Lung cancer clinical decision support demands precise reasoning across complex, multi-stage oncological workflows. Existing multimodal large language models (MLLMs) fail to handle…
OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering
Yifan Zhu, Xinyu Mu, Tao Feng +3
Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA…
SliderQuant: Accurate Post-Training Quantization for LLMs
Shigeng Wang, Chao Li, Yangyuxuan Kang +3
In this paper, we address post-training quantization (PTQ) for large language models (LLMs) from an overlooked perspective: given a pre-trained high-precision LLM, the predominant…
Toward Clinically Explainable AI for Medical Diagnosis: A Foundation Model with Human-Compatible Reasoning via Reinforcement Learning
Qika Lin, Yifan Zhu, Bin Pu +14
The clinical adoption of artificial intelligence (AI) in medical diagnostics is critically hampered by its black-box nature, which prevents clinicians from verifying the rationale…
Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning
Jiangfeng Sun, Sihao He, Zhonghong Ou +1
Multimodal sentiment analysis (MSA) aims to infer emotional states by effectively integrating textual, acoustic, and visual modalities. Despite notable progress, existing multimoda…
TSVC:Tripartite Learning with Semantic Variation Consistency for Robust Image-Text Retrieval
Shuai Lyu, Zijing Tian, Zhonghong Ou +5
Cross-modal retrieval maps data under different modality via semantic relevance. Existing approaches implicitly assume that data pairs are well-aligned and ignore the widely existi…