5 papers
Segment to Focus: Guiding Latent Action Models in the Presence of Distractors
Marcus Fechner, Hamza Adnan, Constantin C. Lüth +3
Latent action models (LAMs) offer a promising path to pre-training embodied agents on large amounts of action-free video. They infer latent actions between consecutive observations…
Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription
Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin +1
Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or r…
SIMA 2: A Generalist Embodied Agent for Virtual Worlds
SIMA team, Adrian Bolton, Alexander Lerchner +63
We introduce SIMA 2, a generalist embodied agent that understands and acts in a wide variety of 3D virtual worlds. Built upon a Gemini foundation model, SIMA 2 represents a signifi…
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
Zheng Xiong, Kang Li, Zilin Wang +3
Built upon language and vision foundation models with strong generalization ability and trained on large-scale robotic data, Vision-Language-Action (VLA) models have recently emerg…
Imagined Autocurricula
Ahmet H. Güzel, Matthew Thomas Jackson, Jarek Luca Liesen +4
Training agents to act in embodied environments typically requires vast training data or access to accurate simulation, neither of which exists for many cases in the real world. In…