most citedAltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

11 citations · 19 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV20231 cited

Multimodal Large Language Model for Visual Navigation

Yao-Hung Hubert Tsai, Vansh Dhar, Jialu Li +2

Recent efforts to enable visual navigation using large language models have mainly focused on developing complex prompt systems. These systems incorporate instructions, observation…

cs.CV20233 cited

Mobile V-MoEs: Scaling Down Vision Transformers via Sparse Mixture-of-Experts

Erik Daxberger, Floris Weers, Bowen Zhang +7

Sparse Mixture-of-Experts models (MoEs) have recently gained popularity due to their ability to decouple model size from inference efficiency by only activating a small subset of t…

cs.CV2023

Likelihood-Based Text-to-Image Evaluation with Patch-Level Perceptual and Semantic Credit Assignment

Qi Chen, Chaorui Deng, Zixiong Huang +3

Text-to-image synthesis has made encouraging progress and attracted lots of public attention recently. However, popular evaluation metrics in this area, like the Inception Score an…

cs.CV20231 cited

Category Feature Transformer for Semantic Segmentation

Quan Tang, Chuanjian Liu, Fagui Liu +5

Aggregation of multi-stage features has been revealed to play a significant role in semantic segmentation. Unlike previous methods employing point-wise summation or concatenation f…

eess.SP20233 cited

Semantic Communications with Variable-Length Coding for Extended Reality

Bowen Zhang, Zhijin Qin, Geoffrey Ye Li

Wireless extended reality (XR) has attracted wide attentions as a promising technology to improve users' mobility and quality of experience. However, the ultra-high data rate requi…

cs.CV2023

STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

Chen Chen, Bowen Zhang, Liangliang Cao +7

Image and text retrieval is one of the foundational tasks in the vision and language domain with multiple real-world applications. State-of-the-art approaches, e.g. CLIP, ALIGN, re…