55 citations · 118 across the 13 of their papers we have counts for
13 papers · 1 filter
RadarSim: Simulating Single-Chip Radar via Multimodal Neural Fields
Chuhan Chen, Tianshu Huang, Akarsh Prabhakara +5
Radars are an ideal complement to cameras: both are inexpensive, solid-state sensors, with cameras offering fine angular resolution, while radars provide metric depth and robustnes…
Visual Self-Refinement for Autoregressive Models
Jiamian Wang, Ziqi Zhou, Chaithanya Kumar Mummadi +5
Autoregressive models excel in sequential modeling and have proven to be effective for vision-language data. However, the spatial nature of visual signals conflicts with the sequen…
Text-driven Prompt Generation for Vision-Language Models in Federated Learning
Chen Qiu, Xingyu Li, Chaithanya Kumar Mummadi +4
Prompt learning for vision-language models, e.g., CoOp, has shown great success in adapting CLIP to different downstream tasks, making it a promising solution for federated learnin…
Zero-Shot Visual Classification with Guided Cropping
Piyapat Saranrittichai, Mauricio Munoz, Volker Fischer +1
Pretrained vision-language models, such as CLIP, show promising zero-shot performance across a wide variety of datasets. For closed-set classification tasks, however, there is an i…
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
Jan Hendrik Metzen, Piyapat Saranrittichai, Chaithanya Kumar Mummadi
Classifiers built upon vision-language models such as CLIP have shown remarkable zero-shot performance across a broad range of image classification tasks. Prior work has studied di…
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
Bang An, Sicheng Zhu, Michael-Andrei Panaitescu-Liess +2
Vision-language models like CLIP are widely used in zero-shot image classification due to their ability to understand various visual concepts and natural language descriptions. How…