2 papers
cs.CV2025
Point Cloud Quantization through Multimodal Prompting for 3D Understanding
Hongxuan Li, Wencheng Zhu, Huiying Xu +2
Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiven…
cs.CV2025
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
Wencheng Zhu, Yuexin Wang, Hongxuan Li +2
Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing approaches primarily rely on parameter-effici…