10 papers
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
CurveShift: Is Agent Progress Scalar? Separating Level from Shape
Hanwen Xing, Pengyun Wang, BingXu Meng +8
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +22
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchma…
Disparities In Negation Understanding Across Languages In Vision-Language Models
Charikleia Moraitaki, Sarah Pan, Skyler Pulling +3
Vision-language models (VLMs) exhibit affirmation bias: a systematic tendency to select positive captions ("X is present") even when the correct description contains negation ("no…
SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
Sepehr Kazemi Ranjbar, Kumail Alhamoud, Marzyeh Ghassemi
Vision-Language Models (VLMs) struggle with negation. Given a prompt like "retrieve (or generate) a street scene without pedestrians," they often fail to respect the "not." Existin…
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
Yubin Kim, Hyewon Jeong, Shan Chen +24
Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence an…