Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve
Dan C. Hsu, Luke Lu
Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success cri…
cs.AI2026
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
Dan C. Hsu, Luke Lu
Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is…