23 citations · 26 across the 4 of their papers we have counts for
4 papers · 1 filter
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
Yin Xie, Kaicheng Yang, Peirou Liang +7
Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to…
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
Qian Zhang, Xiangzi Dai, Ninghua Yang +3
VAR is a new generation paradigm that employs 'next-scale prediction' as opposed to 'next-token prediction'. This innovative transformation enables auto-regressive (AR) transformer…
Point-Voxel Adaptive Feature Abstraction for Robust Point Cloud Classification
Lifa Zhu, Changwei Lin, Chen Zheng +1
Great progress has been made in point cloud classification with learning-based methods. However, complex scene and sensor inaccuracy in real-world application make point cloud data…
Point Cloud Registration using Representative Overlapping Points
Lifa Zhu, Dongrui Liu, Changwei Lin +4
3D point cloud registration is a fundamental task in robotics and computer vision. Recently, many learning-based point cloud registration methods based on correspondences have emer…