Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Zenghui Zhou, Xiaoyang Li, Xiaoxuan Qiao +2
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device tr…
cs.AI2026
MemTX: Transactional Belief Commit for Stateful Agent Memory
Xiaoyang Li, Yiqi Wang, Haohui Lu +5
LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with real side effects. Current a…
cs.AI2026
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
Hao Jiang, Gangtao Xin, Yingdi Huang +35
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-sc…