activity
20192026
most citedQ-ViT: Accurate and Fully Quantized Low-bit Vision Transformer

31 citations · 74 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

GenOpticalFlow: A Generative Approach to Unsupervised Optical Flow Learning

Yixuan Luo, Feng Qiao, Zhexiao Xiong +2

Optical flow estimation is a fundamental problem in computer vision, yet the reliance on expensive ground-truth annotations limits the scalability of supervised approaches. Althoug…

cs.CV2024

CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset

Zhiming Wang, Mingze Wang, Sheng Xu +2

Remote Sensing Image Change Captioning (RSICC) aims to generate natural language descriptions of surface changes between multi-temporal remote sensing images, detailing the categor…

cs.CV20241 cited

P4Q: Learning to Prompt for Quantization in Visual-language Models

Huixin Sun, Runqi Wang, Yanjing Li +4

Large-scale pre-trained Vision-Language Models (VLMs) have gained prominence in various visual and multimodal tasks, yet the deployment of VLMs on downstream application platforms…

cs.CV20231 cited

Representation Disparity-aware Distillation for 3D Object Detection

Yanjing Li, Sheng Xu, Mingbao Lin +3

In this paper, we focus on developing knowledge distillation (KD) for compact 3D detectors. We observe that off-the-shelf KD methods manifest their efficacy only when the teacher m…

cs.CV2023

DCP-NAS: Discrepant Child-Parent Neural Architecture Search for 1-bit CNNs

Yanjing Li, Sheng Xu, Xianbin Cao +4

Neural architecture search (NAS) proves to be among the effective approaches for many tasks by generating an application-adaptive neural architecture, which is still challenged by…

cs.CV20233 cited

Bi-ViT: Pushing the Limit of Vision Transformer Quantization

Yanjing Li, Sheng Xu, Mingbao Lin +4

Vision transformers (ViTs) quantization offers a promising prospect to facilitate deploying large pre-trained networks on resource-limited devices. Fully-binarized ViTs (Bi-ViT) th…