most citedGPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

10 citations · 21 across the 4 of their papers we have counts for

collaborators

5 papers

cs.IR2024

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders

Yupeng Hou, Jiacheng Li, Xiangjun Fu +4

Feature engineering has long been central to recommender systems, yet effectively leveraging textual item features remains challenging. Recent advances in large language models (LL…

cs.CL20232 cited

MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model Evaluation

Zexue He, Yu Wang, An Yan +5

Curated datasets for healthcare are often limited due to the need of human annotations from experts. In this paper, we present MedEval, a multi-level, multi-task, and multi-domain…

cs.CV202310 cited

GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

An Yan, Zhengyuan Yang, Wanrong Zhu +9

We present MM-Navigator, a GPT-4V-based agent for the smartphone graphical user interface (GUI) navigation task. MM-Navigator can interact with a smartphone screen as human users,…

cs.CV2023

Driving through the Concept Gridlock: Unraveling Explainability Bottlenecks in Automated Driving

Jessica Echterhoff, An Yan, Kyungtae Han +3

Concept bottleneck models have been successfully used for explainable machine learning by encoding information within the model with a set of human-defined concepts. In the context…

cs.CV20239 cited

Robust and Interpretable Medical Image Classifiers via Concept Bottleneck Models

An Yan, Yu Wang, Yiwu Zhong +8

Medical image classification is a critical problem for healthcare, with the potential to alleviate the workload of doctors and facilitate diagnoses of patients. However, two challe…