4 papers · 1 filter
RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection
Zhihao Zhang, Gengwei Zhang, Tianlong Chen +1
Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories…
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Jiancheng Huang, Gengwei Zhang, Zequn Jie +5
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal spa…
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
Siyu Jiao, Gengwei Zhang, Yinlong Qian +6
This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVA…
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
Gengwei Zhang, Liyuan Wang, Guoliang Kang +2
In recent years, continual learning with pre-training (CLPT) has received widespread interest, instead of its traditional focus of training from scratch. The use of strong pre-trai…