works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AI2026

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

Xi Fang, Weijie Xu, Yingqiang Ge +3

The paper introduces DRIFTLENS, a framework for measuring how injecting user-specific memory into personalized language models changes the models' reasoning steps, and evaluates me…

cs.AI2026

The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs

Xi Fang, Weijie Xu, Yuchong Zhang +3

When an AI assistant remembers that Sarah is a single mother working two jobs, does it interpret her stress differently than if she were a wealthy executive? As personalized AI sys…

cs.AI2026

Can MLLMs "Read" What is Missing?

Jindi Guo, Chaozheng Huang, Xi Fang

We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text directly from visual context.…

cs.CL2025

SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions

Weijie Xu, Shixian Cui, Xi Fang +3

Large language models (LLMs) are increasingly evaluated on single-answer multiple-choice tasks, yet many real-world problems require identifying all correct answers from a set of o…

cs.CL2025

Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective

Weijie Xu, Yiwen Wang, Chi Xue +4

Large Language Models (LLMs) often generate responses with inherent biases, undermining their reliability in real-world applications. Existing evaluation methods often overlook bia…