1 paper · 1 filter
Kohei Uehara, Nabarun Goswami, Hanqin Wang +10
The increasing demand for intelligent systems capable of interpreting and reasoning about visual content requires the development of large Vision-and-Language Models (VLMs) that ar…