2 citations · 6 across the 15 of their papers we have counts for
16 papers
Grounded 3D-Aware Spatial Vision-Language Modeling
An-Chieh Cheng, Yang Fu, Yatai Ji +12
We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding-…
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
Dongyun Zou, Zhuoyang Zhang, Junyu Chen +8
We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while…
Hide to Guide: Learning via Semantic Masking
Ruitao Liu, Qinghao Hu, Alex Hu +6
Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but its effectiveness is often limi…
: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Physical Intelligence, Bo Ai, Ali Amin +85
We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language i…
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
Luke J. Huang, Zhuoyang Zhang, Qinghao Hu +2
Asynchronous reinforcement learning has become increasingly central to scaling LLM post-training, delivering major throughput gains by decoupling rollout generation from policy upd…
ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
Zhuoyang Zhang, Shang Yang, Qinghao Hu +5
Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We…