3 papers
cs.LG2026
Miles v0.1: Production-Level Post-Training
RadixArk, :, Tom Chen +11
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-lear…
cs.AI2026
xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
Yongchang Peng, Qingshui Gu, Liya Zhu +31
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world…
cs.CL2026
Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
Kangjia Zhao, Jiajun Li, Haozhan Shen +8
Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass clos…