activity
20222024
most citedFerret: Refer and Ground Anything Anywhere at Any Granularity

43 citations · 45 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2024

Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Haotian Zhang, Haoxuan You, Philipp Dufter +8

While Ferret seamlessly integrates regional understanding into the Large Language Model (LLM) to facilitate its referring and grounding capability, it poses certain limitations: co…

cs.CV202343 cited

Ferret: Refer and Ground Anything Anywhere at Any Granularity

Haoxuan You, Haotian Zhang, Zhe Gan +6

We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding op…

cs.CV20232 cited

On Uniform Scalar Quantization for Learned Image Compression

Haotian Zhang, Li Li, Dong Liu

Learned image compression possesses a unique challenge when incorporating non-differentiable quantization into the gradient-based training of the networks. Several quantization sur…

cs.CV2022

Sobolev Training for Implicit Neural Representations with Approximated Image Derivatives

Wentao Yuan, Qingtian Zhu, Xiangyue Liu +3

Recently, Implicit Neural Representations (INRs) parameterized by neural networks have emerged as a powerful and promising tool to represent different kinds of signals due to its c…

cs.CV2022

Spotting Temporally Precise, Fine-Grained Events in Video

James Hong, Haotian Zhang, Michaël Gharbi +2

We introduce the task of spotting temporally precise, fine-grained events in video (detecting the precise moment in time events occur). Precise spotting requires models to reason g…