10 papers
TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning
Yushen Liang, Yue Peng, Baosheng Jin +6
Vision--Language--Action (VLA) policies such as and OpenVLA perform well on many manipulation tasks, but they are often reactive: the next action is predicted from the c…
GVLA: Geometric inductive bias for Vision-Language-Action Models
Yue Peng, Yongzhe Zhao, Artur Habuda +5
Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their…
Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition
Jinnuo Liu, Yue Peng, Jinhan Niu +1
Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a function name: models must coordin…
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
Ailin Huang, Ang Li, Aobo Kong +213
We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most wh…
STEP3-VL-10B Technical Report
Ailin Huang, Chengyuan Yao, Chunrui Han +90
We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-…
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
Jingcheng Hu, Yinmin Zhang, Shijie Shang +17
We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability…