11 papers
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang +3
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on t…
SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution
Aojie Yuan, Yi Nian, Haiyue Zhang +2
Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary label…
AgentIR: A Workload-Adaptive Cascade Retrieval Substrate for Long-Term Conversational Memory
Aojie Yuan, Haiyue Zhang, Shahin Nazarian
Long-term conversational memory is a retrieval workload classical IR was not built for: the index grows during the query stream, query types shift intra-session, and the latency bu…
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Xiaolin Zhou, Aojie Yuan, Zheng Luo +12
Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user…
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang +2
Chain-of-thought (CoT) prompting assumes that generated reasoning reflects a model's internal computation. We show this assumption is wrong in a specific, measurable way: models in…
Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models
Aojie Yuan, Zhiyuan Su
Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether they share a common interna…