2 papers
cs.CV2025
LLaVA-RE: Binary Image-Text Relevancy Evaluation with Multimodal Large Language Model
Tao Sun, Oliver Liu, JinJin Li +1
Multimodal generative AI usually involves generating image or text responses given inputs in another modality. The evaluation of image-text relevancy is essential for measuring res…
cs.RO2024
What do we learn from a large-scale study of pre-trained visual representations in sim and real environments?
Sneha Silwal, Karmesh Yadav, Tingfan Wu +10
We present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involve…