activity
20242026
most citedEvaluating SAM2 for Video Semantic Segmentation

1 citations · 2 across the 15 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.CV2026

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

Syed Ariff Syed Hesham, Yun Liu, Guolei Sun +4

Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense spatiotemporal tokens whose qu…

cs.CV2026

DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation

Xifeng Xue, Xiaokang Wang, Zirui Li +2

Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as…

cs.CV2026

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models

Yi Zhong, Haotong Qin, Xindong Zhang +2

Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However, existing PTQ methods often deg…

eess.IV2026

DINO-Mix: Distilling Foundational Knowledge with Cross-Domain CutMix for Semi-supervised Class-imbalanced Medical Image Segmentation

Xinyu Liu, Guolei Sun

Semi-supervised learning (SSL) has emerged as a critical paradigm for medical image segmentation, mitigating the immense cost of dense annotations. However, prevailing SSL framewor…

cs.LG2026

Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization

Yuli Zhou, Qingxuan Chen, Luca Benini +2

Adaptive Rounding has emerged as an alternative to round-to-nearest (RTN) for post-training quantization by enabling cross-element error cancellation. Yet, dense and element-wise r…

cs.CV2026

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

Bingwen Zhu, Yuqian Fu, Qiaole Dong +6

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating…