6 papers
MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
Zhiyuan Liu, Tinghong Ye, Chenghao Liu +2
Long-horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this problem by compressing and rewriting the hist…
MOSS-Audio Technical Report
Chen Yang, Chufan Yu, Hanfu Chen +27
MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…
PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
Junnan Nie, Jiayi Li, Jiachen Zhang +5
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an…
VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis
Jiachen Zhang, Junyi Lao, Chenghao Liu +5
Finite Element Analysis (FEA) serves as the cornerstone of modern engineering design. However, its workflow is inherently complex and relies heavily on domain expertise. Although r…
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
Chenghao Liu, Jiachen Zhang, Chengxuan Li +4
Vision-Language-Action (VLA) models process visual inputs independently at each timestep, discarding valuable temporal information inherent in robotic manipulation tasks. This fram…
MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning
Chenghao Liu, Zhimu Zhou, Jiachen Zhang +3
Vision-and-Language Navigation (VLN) requires an agent to interpret natural language instructions and navigate complex environments. Current approaches often adopt a "black-box" pa…