12 citations · 21 across the 4 of their papers we have counts for
7 papers
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Jianwei Yang, Hao Zhang, Feng Li +3
We present Set-of-Mark (SoM), a new visual prompting method, to unleash the visual grounding abilities of large multimodal models (LMMs), such as GPT-4V. As illustrated in Fig. 1 (…
Open-NeRF: Towards Open Vocabulary NeRF Decomposition
Hao Zhang, Fang Li, Narendra Ahuja
In this paper, we address the challenge of decomposing Neural Radiance Fields (NeRF) into objects from an open vocabulary, a critical task for object manipulation in 3D reconstruct…
Diff-Retinex: Rethinking Low-light Image Enhancement with A Generative Diffusion Model
Xunpeng Yi, Han Xu, Hao Zhang +2
In this paper, we rethink the low-light image enhancement task and propose a physically explainable and generative diffusion model for low-light image enhancement, termed as Diff-R…
Long-term Leap Attention, Short-term Periodic Shift for Video Classification
Hao Zhang, Lechao Cheng, Yanbin Hao +1
Video transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes times longer sequence than the latter under the curren…
Parameterization of Cross-Token Relations with Relative Positional Encoding for Vision MLP
Zhicai Wang, Yanbin Hao, Xingyu Gao +4
Vision multi-layer perceptrons (MLPs) have shown promising performance in computer vision tasks, and become the main competitor of CNNs and vision Transformers. They use token-mixi…
Small Object Detection Based on Modified FSSD and Model Compression
Qingcai Wang, Hao Zhang, Xianggong Hong +1
Small objects have relatively low resolution, the unobvious visual features which are difficult to be extracted, so the existing object detection methods cannot effectively detect…