activity
20192023
most citedMultimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models

8 citations · 11 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL2023★ 1 cited

Language Models as Black-Box Optimizers for Vision-Language Models

Shihong Liu, Zhiqiu Lin, Samuel Yu +4

Vision-language models (VLMs) pre-trained on web-scale datasets have demonstrated remarkable capabilities on downstream tasks when fine-tuned with minimal data. However, many VLMs…

cs.CV2023★ 8 cited

Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models

Zhiqiu Lin, Samuel Yu, Zhiyi Kuang +2

The ability to quickly learn a new task with minimal instruction - known as few-shot learning - is a central aspect of intelligent agents. Classical few-shot benchmarks make use of…

cs.LG2022★ 2 cited

PACS: A Dataset for Physical Audiovisual CommonSense Reasoning

Samuel Yu, Peter Wu, Paul Pu Liang +2

In order for AI to be safely deployed in real-world scenarios such as hospitals, schools, and the workplace, it must be able to robustly reason about the physical world. Fundamenta…

cs.CV2019

Street Crossing Aid Using Light-weight CNNs for the Visually Impaired

Samuel Yu, Heon Lee, Jung Hoon Kim

In this paper, we address an issue that the visually impaired commonly face while crossing intersections and propose a solution that takes form as a mobile application. The applica…

cs.CV2019

LYTNet: A Convolutional Neural Network for Real-Time Pedestrian Traffic Lights and Zebra Crossing Recognition for the Visually Impaired

Samuel Yu, Heon Lee, John Kim

Currently, the visually impaired rely on either a sighted human, guide dog, or white cane to safely navigate. However, the training of guide dogs is extremely expensive, and canes…