4 papers
Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models
Xingyu Ding, Yuzhong Zhao, Yang Wu +4
Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-ag…
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
Hongyu Ding, Sizhuo Zhang, Ziming Xu +13
Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dom…
HGCN2SP: Hierarchical Graph Convolutional Network for Two-Stage Stochastic Programming
Yang Wu, Yifan Zhang, Zhenxing Liang +1
Two-stage Stochastic Programming (2SP) is a standard framework for modeling decision-making problems under uncertainty. While numerous methods exist, solving such problems with man…
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs
Yang Wu, Yifan Zhang, Yiwei Wang +5
While Large Language Models (LLMs) demonstrate impressive reasoning capabilities, growing evidence suggests much of their success stems from memorized answer-reasoning patterns rat…