Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Segment to Focus: Guiding Latent Action Models in the Presence of Distractors
Marcus Fechner, Hamza Adnan, Constantin C. Lüth +3
Latent action models (LAMs) offer a promising path to pre-training embodied agents on large amounts of action-free video. They infer latent actions between consecutive observations…
cs.LG2026
Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription
Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin +1
Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or r…