5 papers
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation
Zhuohong Chen, Zhengxian Wu, Zirui Liao +6
Vision-centric retrieval for VQA requires retrieving images to supply missing visual cues and integrating them into the reasoning process. However, selecting the right images and i…
TokenUnify: Scaling Up Autoregressive Pretraining for Neuron Segmentation
Yinda Chen, Haoyuan Shi, Xiaoyu Liu +5
Neuron segmentation from electron microscopy (EM) volumes is crucial for understanding brain circuits, yet the complex neuronal structures in high-resolution EM images present sign…
Dual form Complementary Masking for Domain-Adaptive Image Segmentation
Jiawen Wang, Yinda Chen, Xiaoyu Liu +4
Recent works have correlated Masked Image Modeling (MIM) with consistency regularization in Unsupervised Domain Adaptation (UDA). However, they merely treat masking as a special fo…
CBQ: Cross-Block Quantization for Large Language Models
Xin Ding, Xiaoyu Liu, Zhijun Tu +8
Post-training quantization (PTQ) has played a key role in compressing large language models (LLMs) with ultra-low costs. However, existing PTQ methods only focus on handling the ou…
QMamba: Post-Training Quantization for Vision State Space Models
Yinglong Li, Xiaoyu Liu, Jiacheng Li +3
State Space Models (SSMs), as key components of Mamaba, have gained increasing attention for vision models recently, thanks to their efficient long sequence modeling capability. Gi…