6 papers
Structure-Guided Visual Perturbation Neutralization for LVLMs
Yuanhe Zhang, Xueting Wang, YanBin Ren +6
Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface through which adversarial pert…
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
Xinhan Zheng, Huyu Wu, Xueting Wang +2
Multimodal large language models (MLLMs) exhibit a pronounced preference for textual inputs when processing vision-language data, limiting their ability to reason effectively from…
FMASH: Advancing Traditional Chinese Medicine Formula Recommendation with Efficient Fusion of Multiscale Associations of Symptoms and Herbs
Xinhan Zheng, Xueting Wang, Ruotai Li +4
Traditional Chinese medicine (TCM) exhibits remarkable therapeutic efficacy in healthcare through patient-specific formulas. However, current AI-based TCM formula recommendation mo…
Jingfang: An LLM-Based Multi-Agent System for Precise Medical Consultation and Syndrome Differentiation in Traditional Chinese Medicine
Yehan Yang, Tianhao Ma, Ruotai Li +2
The practice of Traditional Chinese Medicine (TCM) requires profound expertise and extensive clinical experience. While Large Language Models (LLMs) offer significant potential in…
The RoboSense Challenge: Sense Anything, Navigate Anywhere, Adapt Across Platforms
Lingdong Kong, Shaoyuan Xie, Zeying Gong +135
Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under…
When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models
Huyu Wu, Meng Tang, Xinhan Zheng +1
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a diverse range of multimodal tasks. However, these models suffer from a core problem know…