3 papers
cs.CV2024
UniDet: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection
Yubin Wang, Zhikang Zou, Xiaoqing Ye +3
We present UniDet, a brand new framework for unified and universal multi-dataset training on 3D detection, enabling robust performance across diverse domains and generalization…
cs.CV2024
MonoFormer: One Transformer for Both Diffusion and Autoregression
Chuyang Zhao, Yuxing Song, Wenhao Wang +5
Most existing multimodality methods use separate backbones for autoregression-based discrete text generation and diffusion-based continuous visual generation, or the same backbone…
cs.CV2024
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
Jing Hao, Yuxiang Zhao, Song Chen +6
Multimodal Large Language Models (MLLMs) have shown promise in a broad range of vision-language tasks with their strong reasoning and generalization capabilities. However, they hea…