activity
20242026
most citedEfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss

2 citations · 6 across the 15 of their papers we have counts for

collaborators

16 papers

cs.CV2026

Grounded 3D-Aware Spatial Vision-Language Modeling

An-Chieh Cheng, Yang Fu, Yatai Ji +12

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding-…

cs.CV2026

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

Dongyun Zou, Zhuoyang Zhang, Junyu Chen +8

We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while…

cs.LG2026

Hide to Guide: Learning via Semantic Masking

Ruitao Liu, Qinghao Hu, Alex Hu +6

Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but its effectiveness is often limi…

cs.LG2026

: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Physical Intelligence, Bo Ai, Ali Amin +85

We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language i…

cs.LG2026

Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs

Luke J. Huang, Zhuoyang Zhang, Qinghao Hu +2

Asynchronous reinforcement learning has become increasingly central to scaling LLM post-training, delivering major throughput gains by decoupling rollout generation from policy upd…

cs.RO2026

ForeAct: Steering Your VLA with Efficient Visual Foresight Planning

Zhuoyang Zhang, Shang Yang, Qinghao Hu +5

Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We…