3 papers
cs.CV2025
Visual Generation Tuning
Jiahao Guo, Sinan Du, Jingfeng Yao +7
Large Vision Language Models (VLMs) effectively bridge the modality gap through extensive pretraining, acquiring sophisticated visual representations aligned with language. However…
cs.CV2025
VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
Sinan Du, Jiahao Guo, Bo Li +8
Unifying multimodal understanding, generation and reconstruction representation in a single tokenizer remains a key challenge in building unified models. Previous research predomin…
cs.CV2024
XS-VID: An Extremely Small Video Object Detection Dataset
Jiahao Guo, Ziyang Xu, Lianjun Wu +3
Small Video Object Detection (SVOD) is a crucial subfield in modern computer vision, essential for early object discovery and detection. However, existing SVOD datasets are scarce…