From the 1 of 6 linked papers with an AI index.
6 papers
RL-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer +4
The paper presents RL², an adaptive test‑time steering framework that uses offline reinforcement learning on latent features from a frozen Vision‑Language‑Action model to compose a…
HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Tasks
Nandiraju Gireesh, Yuanliang Ju, Chaoyi Xu +3
Recent advances in generative models have shown promise in generating behavior plans for long-horizon, sparse reward tasks. While these approaches have achieved promising results,…
Adaptive Q-Chunking for Offline-to-Online Reinforcement Learning
Nandiraju Gireesh, Yuanliang Ju, He Wang
Offline-to-online reinforcement learning with action chunking eliminates multi-step off-policy bias and enables temporally coherent exploration, but all existing methods use a fixe…
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
Yuanchen Ju, Yongyuan Liang, Yen-Jen Wang +7
Mobile manipulators in households must both navigate and manipulate. This requires a compact, semantically rich scene representation that captures where objects are, how they funct…
SAFE: Multitask Failure Detection for Vision-Language-Action Models
Qiao Gu, Yuanliang Ju, Shengxiang Sun +4
While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on no…
ImOV3D: Learning Open-Vocabulary Point Clouds 3D Object Detection from Only 2D Images
Timing Yang, Yuanliang Ju, Li Yi
Open-vocabulary 3D object detection (OV-3Det) aims to generalize beyond the limited number of base categories labeled during the training phase. The biggest bottleneck is the scarc…