3 papers
cs.CV2025
Point Cloud Quantization through Multimodal Prompting for 3D Understanding
Hongxuan Li, Wencheng Zhu, Huiying Xu +2
Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiven…
cs.CV2025
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
Wencheng Zhu, Yuexin Wang, Hongxuan Li +2
Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing approaches primarily rely on parameter-effici…
cs.CV2025
CKD: Contrastive Knowledge Distillation from A Sample-wise Perspective
Wencheng Zhu, Xin Zhou, Pengfei Zhu +2
In this paper, we propose a simple yet effective contrastive knowledge distillation framework that achieves sample-wise logit alignment while preserving semantic consistency. Conve…