3 papers
cs.AI2026
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI
Sean Wu, Pan Lu, Yupeng Chen +7
AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances…
cs.CL2024
Evaluating Spatial Understanding of Large Language Models
Yutaro Yamada, Yihan Bao, Andrew K. Lampinen +2
Large language models (LLMs) show remarkable capabilities across a variety of tasks. Despite the models only seeing text in training, several recent studies suggest that LLM repres…
cs.CV2024
When are Lemons Purple? The Concept Association Bias of Vision-Language Models
Yutaro Yamada, Yingtian Tang, Yoyo Zhang +1
Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-to-text retrieval. However, such performance does not…