135 citations · 853 across the 69 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 4 cited
Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants
Tianyu Yu, Jinyi Hu, Yuan Yao +10
Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial…
cs.CV2023★ 10 cited
Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models
Chi Chen, Ruoyu Qin, Fuwen Luo +4
Recently, Multimodal Large Language Models (MLLMs) that enable Large Language Models (LLMs) to interpret images through visual instruction tuning have achieved significant success.…
cs.CV2021
Visual Distant Supervision for Scene Graph Generation
Yuan Yao, Ao Zhang, Xu Han +5
Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer visi…