From the 1 of 3 linked papers with an AI index.
3 papers
cs.RO2026
Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection
Yi Wang, Wendi Chen, Zimo Wen +8
The paper introduces LIFT, a post‑training method that adds reactive force feedback to pretrained vision‑language‑action policies, enabling them to handle contact‑rich manipulation…
cs.CV2025
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Jinheng Xie, Weijia Mao, Zechen Bai +7
We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discret…
cs.CV2025
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Kepan Nan, Rui Xie, Penghao Zhou +6
Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challeng…