8 papers
Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks
Boxiu Li, Zimo Wen, Yijia Fan +24
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a…
Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization
Chenghao Liu, Yu Zhang, Zhongtao Jiang +7
Embedding-based retrieval ranks items by their similarity to a query in a shared vector space and usually aims to return the highest-scoring items. In many production settings this…
PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
Junnan Nie, Jiayi Li, Jiachen Zhang +5
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an…
VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis
Jiachen Zhang, Junyi Lao, Chenghao Liu +5
Finite Element Analysis (FEA) serves as the cornerstone of modern engineering design. However, its workflow is inherently complex and relies heavily on domain expertise. Although r…
What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies
Jiachen Zhang, Junnan Nie, Junyi Lao +4
Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Their frozen representations nev…
DIRCR: Dual-Inference Rule-Contrastive Reasoning for Solving RAVENs
Jiachen Zhang, Chengtai Li, Jianfeng Ren +3
Abstract visual reasoning remains challenging as existing methods often prioritize either global context or local row-wise relations, failing to integrate both, and lack intermedia…