1 paper
Yufei Zhan, Yousong Zhu, Zhiyang Chen +3
Replicating the innate human ability to detect all objects based on free-form texts at any granularity remains a formidable challenge for Large Vision Language Models (LVLMs). Curr…