activity
20242026
collaborators

7 papers

cs.CV2026

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

Shenghong Yi, Lin Zhang, Muzian Li +6

Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-…

cs.CV2026

FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding

Kangcong Li, Peng Ye, Lin Zhang +3

Transitioning Multimodal Large Language Models (MLLMs) from offline to online streaming video understanding is essential for continuous perception. However, existing methods lack f…

cs.CV2025

Understanding Cross Task Generalization in Handwriting-Based Alzheimer's Screening via Vision Language Adaptation

Changqing Gong, Huafeng Qin, Mounim A. El-Yacoubi

Alzheimer's disease is a prevalent neurodegenerative disorder for which early detection is critical. Handwriting-often disrupted in prodromal AD-provides a non-invasive and cost-ef…

cs.LG2025

A Survey on Mixup Augmentations and Beyond

Xin Jin, Hongyu Zhu, Siyuan Li +6

As Deep Neural Networks have achieved thrilling breakthroughs in the past decade, data augmentations have garnered increasing attention as regularization techniques when massive la…

cs.CV2025

EM-DARTS: Hierarchical Differentiable Architecture Search for Eye Movement Recognition

Huafeng Qin, Hongyu Zhu, Xin Jin +3

Eye movement biometrics has received increasing attention thanks to its highly secure identification. Although deep learning (DL) models have shown success in eye movement recognit…

cs.CV2024

Hybrid Transformer for Early Alzheimer's Detection: Integration of Handwriting-Based 2D Images and 1D Signal Features

Changqing Gong, Huafeng Qin, Mounîm A. El-Yacoubi

Alzheimer's Disease (AD) is a prevalent neurodegenerative condition where early detection is vital. Handwriting, often affected early in AD, offers a non-invasive and cost-effectiv…