3 papers
cs.AI2026
Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve
Dan C. Hsu, Luke Lu
Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success cri…
cs.AI2026
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
Dan C. Hsu, Luke Lu
Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is…
cs.LG2026
Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback
Ved Sriraman, Peihan Liu, Daniel Hsu +1
Imitation Learning is a natural framework for learning in sequential decision-making systems and has emerged as the dominant paradigm through which we understand language model tra…