6 papers
ViTexSZ: Heterogeneous Vision-Text Knowledge Distillation for EEG Seizure Detection
Chenxi Liu, Mingzhao Li, Yicong Liu +4
Automated seizure detection from electroencephalography (EEG) is essential for continuous neurological monitoring, particularly for subclinical epileptic seizures that may exhibit…
PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
Jianming Liu, Ren Zhu, Jian Xu +4
Solving Partial Differential Equations (PDEs) is a cornerstone of engineering and scientific research. Traditional methods for PDE solving are cumbersome, relying on manual setup a…
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
Shixiong Xu, Chenghao Zhang, Lubin Fan +5
Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained s…
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
Lingdong Shen, Qigqi, Kun Ding +2
Scientific Literature charts often contain complex visual elements, including multi-plot figures, flowcharts, structural diagrams and etc. Evaluating multimodal models using these…
A Survey of Low-shot Vision-Language Model Adaptation via Representer Theorem
Kun Ding, Ying Wang, Gaofeng Meng +1
The advent of pre-trained vision-language foundation models has revolutionized the field of zero/few-shot (i.e., low-shot) image recognition. The key challenge to address under the…
Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation
Kun Ding, Qiang Yu, Haojian Zhang +2
Cache-based approaches stand out as both effective and efficient for adapting vision-language models (VLMs). Nonetheless, the existing cache model overlooks three crucial aspects.…