3 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Keen You, Haotian Zhang, Eldon Schoop +5
Recent advancements in multimodal large language models (MLLMs) have been noteworthy, yet, these general-domain MLLMs often fall short in their ability to comprehend and interact e…
cs.HC2023★ 3 cited
ILuvUI: Instruction-tuned LangUage-Vision modeling of UIs from Machine Conversations
Yue Jiang, Eldon Schoop, Amanda Swearngin +1
Multimodal Vision-Language Models (VLMs) enable powerful applications from their fused understanding of images and language, but many perform poorly on UI tasks due to the lack of…
cs.HC2023
Never-ending Learning of User Interfaces
Jason Wu, Rebecca Krosnick, Eldon Schoop +3
Machine learning models have been trained to predict semantic information about user interfaces (UIs) to make apps more accessible, easier to test, and to automate. Currently, most…