From the 1 of 28 linked papers with an AI index.
10 papers · 1 filter
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
Yanlong Chen, Amirhossein Habibian, Luca Benini +1
Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its pot…
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
Yujie Chen, Haotong Qin, Zhang Zhang +3
State-Space Models (SSMs) have attracted considerable attention in Image Restoration (IR) due to their ability to scale linearly sequence length while effectively capturing long-di…
CamSAM2: Segment Anything Accurately in Camouflaged Videos
Yuli Zhou, Yawei Li, Yuqian Fu +3
Video camouflaged object segmentation (VCOS), aiming at segmenting camouflaged objects that seamlessly blend into their environment, is a fundamental vision task with various real-…
WaveFormer: A Lightweight Transformer Model for sEMG-based Gesture Recognition
Yanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa +3
Human-machine interaction, particularly in prosthetic and robotic control, has seen progress with gesture recognition via surface electromyographic (sEMG) signals.However, classify…
HiM2SAM: Enhancing SAM2 with Hierarchical Motion Estimation and Memory Optimization towards Long-term Tracking
Ruixiang Chen, Guolei Sun, Yawei Li +2
This paper presents enhancements to the SAM2 framework for video object tracking task, addressing challenges such as occlusions, background clutter, and target reappearance. We int…
FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning
Hang Guo, Yawei Li, Taolin Zhang +4
Visual Autoregressive (VAR) modeling has gained popularity for its shift towards next-scale prediction. However, existing VAR paradigms process the entire token map at each scale s…