4 papers
When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition
Chencheng Zhu, Xiaoyang Li, Taotao Cai
Task arithmetic composes skills by adding weight displacements, and merged models are then judged on benchmark suites. We measure when that composition is functionally additive, an…
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Zenghui Zhou, Xiaoyang Li, Xiaoxuan Qiao +2
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device tr…
MemTX: Transactional Belief Commit for Stateful Agent Memory
Xiaoyang Li, Yiqi Wang, Haohui Lu +5
LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with real side effects. Current a…
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
Hao Jiang, Gangtao Xin, Yingdi Huang +35
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-sc…