4 papers
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
Ailin Huang, Ang Li, Aobo Kong +213
We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most wh…
ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
Huaixiao Tou, Ying Zeng, Yuemeng Li +6
We present ShoppingComp, a challenging real-world benchmark for comprehensively evaluating LLM-powered shopping agents on three core capabilities: precise product retrieval, expert…
Step-DeepResearch Technical Report
Chen Hu, Haikuo Du, Heng Wang +64
As LLMs shift toward autonomous agents, Deep Research has emerged as a pivotal metric. However, existing academic benchmarks like BrowseComp often fail to meet real-world demands f…
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems
Yebo Peng, Zixiang Liu, Yaoming Li +6
Evaluating the mathematical capability of Large Language Models (LLMs) is a critical yet challenging frontier. Existing benchmarks fall short, particularly for proof-centric proble…