17 papers
STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation
Syed Ariff Syed Hesham, Yun Liu, Guolei Sun +4
Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense spatiotemporal tokens whose qu…
DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation
Xifeng Xue, Xiaokang Wang, Zirui Li +2
Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as…
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
Yi Zhong, Haotong Qin, Xindong Zhang +2
Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However, existing PTQ methods often deg…
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
Bingwen Zhu, Yuqian Fu, Qiaole Dong +6
Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating…
DINO-Mix: Distilling Foundational Knowledge with Cross-Domain CutMix for Semi-supervised Class-imbalanced Medical Image Segmentation
Xinyu Liu, Guolei Sun
Semi-supervised learning (SSL) has emerged as a critical paradigm for medical image segmentation, mitigating the immense cost of dense annotations. However, prevailing SSL framewor…
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
Xinyu Liu, Guolei Sun, Cheng Wang +2
High-resolution (HR) medical videos are vital for accurate diagnosis, yet are hard to acquire due to hardware limitations and physiological constraints. Clinically, the collected l…