77 citations · 145 across the 9 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 1 cited
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
Yuhang Zang, Hanlin Goh, Josh Susskind +1
Existing vision-language models exhibit strong generalization on a variety of visual domains and tasks. However, such models mainly perform zero-shot recognition in a closed-set ma…
cs.CV2022
Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents
Yao-Hung Hubert Tsai, Hanlin Goh, Ali Farhadi +1
The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human…