5 papers
Vision Language Models Cannot Reason About Physical Transformation
Dezhi Luo, Yijiang Li, Maijunxian Wang +7
Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
Zory Zhang, Pinyuan Feng, Bingyang Wang +7
Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation st…
Towards Interpretable Visual Decoding with Attention to Brain Representations
Pinyuan Feng, Hossein Adeli, Wenxuan Guo +3
Recent work has demonstrated that complex visual stimuli can be decoded from human brain activity using deep generative models, offering new ways to probe how the brain represents…
TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
Yanshu Li, Jianjiang Yang, Tian Yun +3
Multimodal in-context learning (ICL) has emerged as a key mechanism for harnessing the capabilities of large vision-language models (LVLMs). However, its effectiveness remains high…
Better artificial intelligence does not mean better models of biology
Drew Linsley, Pinyuan Feng, Thomas Serre
Deep neural networks (DNNs) once showed increasing alignment with primate perception and neural responses as they improved on vision benchmarks, raising hopes that advances in AI w…