Showing 2024Show all
3 papers · 1 filter
cs.CV2024★ 1 cited
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Chengyue Wu, Xiaokang Chen, Zhiyu Wu +8
In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both…
cs.CV2024
Adapting LLaMA Decoder to Vision Transformer
Jiahao Wang, Wenqi Shao, Mengzhao Chen +7
This work examines whether decoder-only Transformers such as LLaMA, which were originally designed for large language models (LLMs), can be adapted to the computer vision field. We…
cs.CV2024
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
Hritik Bansal, Po-Nien Kung, P. Jeffrey Brantingham +2
Multimodal event argument role labeling (EARL), a task that assigns a role for each event participant (object) in an image is a complex challenge. It requires reasoning over the en…