2 papers
cs.SE2026
FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration
Weihang Ding, Junfei Zhan, Yueting Li +1
Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 13…
cs.LG2026
Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
Weihang Ding, Junfei Zhan
Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model…