From the 1 of 6 linked papers with an AI index.
6 papers
Query Timing Produces Opposite Positional Biases Between LLMs and Humans
Jasin Cekinmez, Addison J. Wu, Thomas L. Griffiths
Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluation…
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Jasin Cekinmez, Addison J. Wu, Raja Marjieh +1
Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans communicate and imagine. Large language model…
SportD: How do VLMs physically strategize?
Jasin Cekinmez, Addison J. Wu, Haotian Xia +11
The paper introduces SportD, a benchmark that tests whether vision‑language models can choose optimal shoot or pass actions in soccer situations, comparing model choices to a value…
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
Ryo Mitsuhashi, Patrick Chen, Isabelle Tseng +2
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are rarely perfect. Recent theore…
Guess the Unified Model: How Much Can We Recover from Generated Images?
Jasin Cekinmez, Ryo Mitsuhashi, Addison J. Wu +1
With unified model-generated images now widespread online, attributing their model of origin offers a path toward transparency and deeper insight into the characteristic behaviors…
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
Jasin Cekinmez, Omid Ghahroodi, Saad Fowad Chandle +2
We introduce ADAM (A Diverse Archive of Mankind), a framework for evaluating and improving multimodal large language models (MLLMs) in biographical reasoning. To the best of our kn…