3 papers
cs.AI2026
SKILL.nb: Selective Formalization and Gated Execution for Durable Agent Workflows
Amine El Hattami, Nicolas Chapados, Christopher Pal
AI agents increasingly turn past experience into reusable artifacts such as code, workflows, and procedural memories. Reuse can improve efficiency, but it also creates a lifecycle…
cs.CE2026
Training Diffusion Language Models for Black-Box Optimization
Zipeng Sun, Can Chen, Ye Yuan +4
We study offline black-box optimization (BBO), aiming to discover improved designs from an offline dataset of designs and labels, a problem common in robotics and DNA with limited…
cs.CL2026
DRBench: A Realistic Benchmark for Enterprise Deep Research
Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol +11
We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions…