2 papers
cs.CV2025
JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse
Muyao Li, Zihao Wang, Kaichen He +2
Recently, action-based decision-making in open-world environments has gained significant attention. Visual Language Action (VLA) models, pretrained on large-scale web datasets, hav…
cs.AI2024
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
Shaofei Cai, Bowei Zhang, Zihao Wang +4
Developing agents that can follow multimodal instructions remains a fundamental challenge in robotics and AI. Although large-scale pre-training on unlabeled datasets (no language i…