4 papers · 1 filter
RadarSim: Simulating Single-Chip Radar via Multimodal Neural Fields
Chuhan Chen, Tianshu Huang, Akarsh Prabhakara +5
Radars are an ideal complement to cameras: both are inexpensive, solid-state sensors, with cameras offering fine angular resolution, while radars provide metric depth and robustnes…
Visual Self-Refinement for Autoregressive Models
Jiamian Wang, Ziqi Zhou, Chaithanya Kumar Mummadi +5
Autoregressive models excel in sequential modeling and have proven to be effective for vision-language data. However, the spatial nature of visual signals conflicts with the sequen…
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
Jan Hendrik Metzen, Piyapat Saranrittichai, Chaithanya Kumar Mummadi
Classifiers built upon vision-language models such as CLIP have shown remarkable zero-shot performance across a broad range of image classification tasks. Prior work has studied di…
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
Bang An, Sicheng Zhu, Michael-Andrei Panaitescu-Liess +2
Vision-language models like CLIP are widely used in zero-shot image classification due to their ability to understand various visual concepts and natural language descriptions. How…