From the 1 of 9 linked papers with an AI index.
9 papers
Query Timing Produces Opposite Positional Biases Between LLMs and Humans
Jasin Cekinmez, Addison J. Wu, Thomas L. Griffiths
Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluation…
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Jasin Cekinmez, Addison J. Wu, Raja Marjieh +1
Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans communicate and imagine. Large language model…
SportD: How do VLMs physically strategize?
Jasin Cekinmez, Addison J. Wu, Haotian Xia +11
The paper introduces SportD, a benchmark that tests whether vision‑language models can choose optimal shoot or pass actions in soccer situations, comparing model choices to a value…
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
Addison J. Wu, Ryan Liu, Xuechunzi Bai +1
As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In t…
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
Ryo Mitsuhashi, Patrick Chen, Isabelle Tseng +2
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are rarely perfect. Recent theore…
Guess the Unified Model: How Much Can We Recover from Generated Images?
Jasin Cekinmez, Ryo Mitsuhashi, Addison J. Wu +1
With unified model-generated images now widespread online, attributing their model of origin offers a path toward transparency and deeper insight into the characteristic behaviors…