6 papers
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation
Zhuohong Chen, Zhengxian Wu, Zirui Liao +6
Vision-centric retrieval for VQA requires retrieving images to supply missing visual cues and integrating them into the reasoning process. However, selecting the right images and i…
Dual form Complementary Masking for Domain-Adaptive Image Segmentation
Jiawen Wang, Yinda Chen, Xiaoyu Liu +4
Recent works have correlated Masked Image Modeling (MIM) with consistency regularization in Unsupervised Domain Adaptation (UDA). However, they merely treat masking as a special fo…
QMamba: Post-Training Quantization for Vision State Space Models
Yinglong Li, Xiaoyu Liu, Jiacheng Li +3
State Space Models (SSMs), as key components of Mamaba, have gained increasing attention for vision models recently, thanks to their efficient long sequence modeling capability. Gi…
Multi-Granularity Semantic Revision for Large Language Model Distillation
Xiaoyu Liu, Yun Zhang, Wei Li +7
Knowledge distillation plays a key role in compressing the Large Language Models (LLMs), which boosts a small-size student model under large teacher models' guidance. However, exis…
UniCompress: Enhancing Multi-Data Medical Image Compression with Knowledge Distillation
Runzhao Yang, Yinda Chen, Zhihong Zhang +6
In the field of medical image compression, Implicit Neural Representation (INR) networks have shown remarkable versatility due to their flexible compression ratios, yet they are co…
TokenUnify: Scaling Up Autoregressive Pretraining for Neuron Segmentation
Yinda Chen, Haoyuan Shi, Xiaoyu Liu +5
Neuron segmentation from electron microscopy (EM) volumes is crucial for understanding brain circuits, yet the complex neuronal structures in high-resolution EM images present sign…