5 papers
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
Tingyang Chen, Shuo Lu, Kang Zhao +11
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet to…
RecThinker: An Agentic Framework for Tool-Augmented Reasoning in Recommendation
Haobo Zhang, Yutao Zhu, Kelong Mao +2
Large Language Models (LLMs) have revolutionized recommendation agents by providing superior reasoning and flexible decision-making capabilities. However, existing methods mainly f…
ChatShopBuddy: Towards Reliable Conversational Shopping Agents via Reinforcement Learning
Yiruo Cheng, Kelong Mao, Tianhao Li +3
Conversational shopping agents represent a critical consumer-facing application of Large Language Model (LLM)-powered agents, yet how to effectively apply post-training Reinforceme…
AI Benchmark Democratization and Carpentry
Gregor von Laszewski, Wesley Brewer, Jeyan Thiyagalingam +28
Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring d…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…