3 papers
cs.LG2026
FlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning
Takanori Yoshimoto, Yang Hu, Naruya Kondo +1
Latent actions provide a compact interface between action-free video and downstream decision-making, yet existing Latent Action Models (LAMs) force every transition through a fixed…
cs.RO2025
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni +32
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a centra…
cs.CV2023
StreamDiffusion: A Pipeline-level Solution for Real-time Interactive Generation
Akio Kodaira, Chenfeng Xu, Toshiki Hazama +8
We introduce StreamDiffusion, a real-time diffusion pipeline designed for interactive image generation. Existing diffusion models are adept at creating images from text or image pr…