11 papers
A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation
Yulin Xue, Siqi Ouyang, Lei Li
Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech ra…
RASST: Retrieval-Augmented Simultaneous Speech Translation
Jiaxuan Luo, Siqi Ouyang, Jiaxing Xu +1
Simultaneous speech translation produces target text incrementally from partial speech input. Recent speech large language models have markedly improved SST quality but still strug…
Bernini: Latent Semantic Planning for Video Diffusion
Bernini Team, Chenchen Liu, Junyi Chen +9
Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimodal inputs with strong seman…
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
Yuxiang Huang, Nuno M. T. Gonçalves, Federico Alvetreti +5
Current hierarchical attention methods, such as NSA and InfLLMv2, select the top-k relevant key-value (KV) blocks based on coarse attention scores and subsequently apply fine-grain…
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
Morunliu Yang, Ruotao Xu, Le Li +6
Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences intr…
Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
Siqi Ouyang, Shuoyang Ding, Oleksii Hrinchuk +4
Simultaneous speech translation (SST) generates translations while receiving partial speech input. Recent advances show that large language models (LLMs) can substantially improve…