activity
20172025
most citedLlama 2: Open Foundation and Fine-Tuned Chat Models

2.7k citations · 4k across the 34 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

Zhiyang Xu, Jiuhai Chen, Zhaojiang Lin +10

Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite thes…

cs.CV2024

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CV2023★ 8 cited

MMViT: Multiscale Multiview Vision Transformers

Yuchen Liu, Natasha Ong, Kaiyan Peng +8

We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…

cs.CV2023

SVT: Supertoken Video Transformer for Efficient Video Understanding

Chenbin Pan, Rui Hou, Hanchao Yu +3

Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video conte…

cs.CV2022★ 7 cited

A Fistful of Words: Learning Transferable Visual Models from Bag-of-Words Supervision

Ajinkya Tejankar, Maziar Sanjabi, Bichen Wu +4

Using natural language as a supervision for training visual recognition models holds great promise. Recent works have shown that if such supervision is used in the form of alignmen…