1 paper
Zachary Huemann, Samuel Church, Joshua D. Warner +7
Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potential applications in enhanced radi…