2 citations · 6 across the 19 of their papers we have counts for
32 papers
Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation
Zhilong Zhang, Wenyu Luo, Haonan Wang +9
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, enabling robots to perform tasks based on natural language instructions and curre…
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models
Zhilong Zhang, Haoxiang Ren, Yihao Sun +6
Vision-Language-Action (VLA) models show strong generalization for robotic control, but finetuning them with reinforcement learning (RL) is constrained by the high cost and safety…
Transformers Are Born Biased: Structural Inductive Biases at Random Initialization and Their Practical Consequences
Siquan Li, Yao Tong, Haonan Wang +1
Transformers underpin modern large language models (LLMs) and are commonly assumed to be behaviorally unstructured at random initialization, with all meaningful preferences emergin…
Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
Haonan Wang, Chao Du, Kenji Kawaguchi +1
Majority voting has proven effective for close-ended question answering by aggregating parallel reasoning traces. However, it is not directly applicable to open-ended reasoning, su…
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
Haonan Wang, Hanyu Zhou, Haoyue Liu +1
We investigate a challenging task of dynamic scene geometry estimation, which requires representing both spatial and temporal features. Typically, existing methods align the two fe…
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
MiroMind Team, Song Bai, Lidong Bing +52
We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike previous agents that only scale…