3 papers
cs.CV2024
VILA: VILA Augmented VILA
Yunhao Fang, Ligeng Zhu, Yao Lu +7
While visual language model architectures and training infrastructures advance rapidly, data curation remains under-explored where quantity and quality become a bottleneck. Existin…
cs.CV2024
Language-Image Models with 3D Understanding
Jang Hyun Cho, Boris Ivanovic, Yulong Cao +8
Multi-modal large language models (MLLMs) have shown incredible capabilities in a variety of 2D vision and language tasks. We extend MLLMs' perceptual capabilities to ground and re…
cs.CV2023
Language-conditioned Detection Transformer
Jang Hyun Cho, Philipp Krähenbühl
We present a new open-vocabulary detection framework. Our framework uses both image-level labels and detailed detection annotations when available. Our framework proceeds in three…