activity
20142023
most citedDetecting Sarcasm in Multimodal Social Platforms

236 citations · 284 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV202343 cited

Ferret: Refer and Ground Anything Anywhere at Any Granularity

Haoxuan You, Haotian Zhang, Zhe Gan +6

We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding op…

cs.CV20232 cited

Efficient-3DiM: Learning a Generalizable Single-image Novel-view Synthesizer in One Day

Yifan Jiang, Hao Tang, Jen-Hao Rick Chang +3

The task of novel view synthesis aims to generate unseen perspectives of an object or scene from a limited set of input images. Nevertheless, synthesizing novel views from a single…

cs.CL2023

Instruction-Following Speech Recognition

Cheng-I Jeff Lai, Zhiyun Lu, Liangliang Cao +1

Conventional end-to-end Automatic Speech Recognition (ASR) models primarily focus on exact transcription tasks, lacking flexibility for nuanced user interactions. With the advent o…

cs.CV20231 cited

RoomDreamer: Text-Driven 3D Indoor Scene Synthesis with Coherent Geometry and Texture

Liangchen Song, Liangliang Cao, Hongyu Xu +4

The techniques for 3D indoor scene capturing are widely used, but the meshes produced leave much to be desired. In this paper, we propose "RoomDreamer", which leverages powerful na…

cs.CV20232 cited

Less is More: Removing Text-regions Improves CLIP Training Efficiency and Robustness

Liangliang Cao, Bowen Zhang, Chen Chen +5

The CLIP (Contrastive Language-Image Pre-training) model and its variants are becoming the de facto backbone in many applications. However, training a CLIP model from hundreds of m…

cs.CV2023

STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

Chen Chen, Bowen Zhang, Liangliang Cao +7

Image and text retrieval is one of the foundational tasks in the vision and language domain with multiple real-world applications. State-of-the-art approaches, e.g. CLIP, ALIGN, re…