2 papers
cs.CL2026
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Zhuochun Li, Youngmin Ko, Ali Keramati +9
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential…
cs.LG2025
LAMP: Extracting Local Decision Surfaces From Large Language Models
Ryan Chen, Youngmin Ko, Zeyu Zhang +5
We introduce LAMP (Local Attribution Mapping Probe), a method that shines light onto a black-box language model's decision surface and studies how reliably a model maps its stated…