collaborators

10 papers

cs.AI2026

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Xiaomin Li, Yuexing Hao, Jianheng Hou +90

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…

cs.CL2026

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Hanwen Xing, Pengyun Wang, BingXu Meng +8

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…

cs.LG2026

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +22

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchma…

cs.CL2026

Disparities In Negation Understanding Across Languages In Vision-Language Models

Charikleia Moraitaki, Sarah Pan, Skyler Pulling +3

Vision-language models (VLMs) exhibit affirmation bias: a systematic tendency to select positive captions ("X is present") even when the correct description contains negation ("no…

cs.CV2025

SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models

Sepehr Kazemi Ranjbar, Kumail Alhamoud, Marzyeh Ghassemi

Vision-Language Models (VLMs) struggle with negation. Given a prompt like "retrieve (or generate) a street scene without pedestrians," they often fail to respect the "not." Existin…

cs.CL2025

Medical Hallucinations in Foundation Models and Their Impact on Healthcare

Yubin Kim, Hyewon Jeong, Shan Chen +24

Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence an…