7 papers · 1 filter
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
Chong Cheng, Peilin Tao, Nanjie Yao +9
Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or…
MoASE++: Mixture of Activation Sparsity Experts with Domain-Adaptive On-policy Distillation for Continual Test Time Adaptation
Ronyu Zhang, Aosong Cheng, Gaole Dai +8
Continual test-time adaptation adapts a source-pretrained model to non-stationary, unlabeled target streams while retaining past competence, yet texture-biased backbones risk error…
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
Hengyu Fang, Yijiang Liu, Yuan Du +2
Vision-Language-Action (VLA) models exhibit unprecedented capabilities for embodied intelligence. However, their extensive computational and memory costs hinder their practical dep…
Fisher-aware Quantization for DETR Detectors with Critical-category Objectives
Huanrui Yang, Yafeng Huang, Zhen Dong +6
The impact of quantization on the overall performance of deep learning models is a well-studied problem. However, understanding and mitigating its effects on a more fine-grained le…
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
Rongyu Zhang, Aosong Cheng, Yulin Luo +8
Continual Test-Time Adaptation (CTTA), which aims to adapt the pre-trained model to ever-evolving target domains, emerges as an important task for vision models. As current vision…
VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
Rongyu Zhang, Zefan Cai, Huanrui Yang +9
Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. However, the conventional finetuning process with randomly sampled data point…