collaborators

14 papers

cs.CV2026

COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

Chenghua Zhu, Zhaolu Kang, Qifan Shi +8

Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sam…

cs.IR2026

Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval

Runhao Li, Xiaoxu Ma, Zhenyu Weng +5

Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual semantic annotation. However, ex…

cs.IR2026

Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing

Runhao Li, Xiaoxu Ma, Zhenyu Weng +5

Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled image-text pairs. However, existi…

cs.LG2026

SPOTR: Spatio-temporal Pooling One-Token Reconstruction for Universal Physiological Signal Self-supervised Learning

Yiyu Gui, Mingzhi Chen, Yuesheng Zhu +2

Physiological signals such as EEG, ECG, and PPG are widely used in clinical monitoring. Recent self-supervised learning (SSL) methods offer an attractive way to leverage unlabeled…

cs.LG2026

MedTS-TTT: Test-Time Training for Medical Time Series Classification

Mingzhi Chen, Yiyu Gui, Guibo Luo

Medical time series (MedTS) signals such as electroencephalography (EEG) and electrocardiography (ECG) support many clinical applications. However, substantial subject-level hetero…

cs.CV2026

DSFedMed: Dual-Scale Federated Medical Image Segmentation via Mutual Distillation Between Foundation and Lightweight Models

Hanwen Zhang, Qiaojin Shen, Yuxi Liu +2

Foundation Models (FMs) have demonstrated strong generalization across diverse vision tasks. However, their deployment in federated settings is hindered by high computational deman…