From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
Xi Fang, Weijie Xu, Yingqiang Ge +3
The paper introduces DRIFTLENS, a framework for measuring how injecting user-specific memory into personalized language models changes the models' reasoning steps, and evaluates me…
cs.AI2026
Stop Comparing LLM Agents Without Disclosing the Harness
Yunbei Zhang, Janet Wang, Yingqiang Ge +3
This position paper argues that, for long-horizon tasks evaluated across models with comparable frontier capability, the agent execution harness, namely the infrastructure layer th…