2 papers
cs.CV2025
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning
Rui Zhao, Qirui Yuan, Jinyu Li +4
End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal La…
cs.CL2024
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
Rui Zhao, Jinyu Li, Ruchao Fan +1
Models for streaming speech translation (ST) can achieve high accuracy and low latency if they're developed with vast amounts of paired audio in the source language and written tex…