3 papers
cs.AI2026
Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents
Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt +5
Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon e…
cs.CV2026
Steerable Visual Representations
Jona Ruthardt, Manu Gaur, Deva Ramanan +2
Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification,…
cs.CL2026
Better Language Models Exhibit Higher Visual Alignment
Jona Ruthardt, Gertjan J. Burghouts, Serge Belongie +1
How well do text-only large language models (LLMs) align with the visual world? We present a systematic evaluation of this question by incorporating frozen representations of vario…