11 papers
HumanCLAW: Can Vision-Language Models Act Through a Body?
Siyao Li, Li Siyao, Jiawei Gu +16
The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
Zican Hu, Xuyang Hu, Yiming Liu +10
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learni…
3D-CoS: A New 3D Reconstruction Paradigm Based on VLM Code Synthesis
Yuhao Wang, Puyi Wang, Linjie Li +3
Most recent 3D reconstruction and editing systems operate on implicit and explicit representations such as NeRF, point clouds, or meshes. While these representations enable high-fi…
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Jiaqi Liu, Shi Qiu, Mairui Li +33
Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple perspectives, experiments fail…
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects
Puyi Wang, Yuhao Wang, Linjie Li +4
Indoor scene synthesis underpins embodied AI, robotic manipulation, and simulation-based policy evaluation, where a useful scene must specify not only what the environment looks li…
Gym-V: A Unified Vision Environment System for Agentic Vision Research
Fanqing Meng, Lingxiao Du, Jiawei Gu +9
As agentic systems increasingly rely on reinforcement learning from verifiable rewards, standardized ``gym'' infrastructure has become essential for rapid iteration, reproducibilit…