4 papers
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Jasin Cekinmez, Addison J. Wu, Raja Marjieh +1
Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans communicate and imagine. Large language model…
SportD: How do VLMs physically strategize?
Jasin Cekinmez, Addison J. Wu, Haotian Xia +6
Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed…
Query Timing Produces Opposite Positional Biases Between LLMs and Humans
Jasin Cekinmez, Addison J. Wu, Thomas L. Griffiths
Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluation…
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
Jasin Cekinmez, Omid Ghahroodi, Saad Fowad Chandle +2
We introduce ADAM (A Diverse Archive of Mankind), a framework for evaluating and improving multimodal large language models (MLLMs) in biographical reasoning. To the best of our kn…