3 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.CV2025
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
Wenwen Tong, Hewei Guo, Dongchuan Ran +23
We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead…
cs.SD2022
Dynamic Latency for CTC-Based Streaming Automatic Speech Recognition With Emformer
Jingyu Sun, Guiping Zhong, Dinghao Zhou +1
An inferior performance of the streaming automatic speech recognition models versus non-streaming model is frequently seen due to the absence of future context. In order to improve…
cs.SD2022★ 3 cited
Locality Matters: A Locality-Biased Linear Attention for Automatic Speech Recognition
Jingyu Sun, Guiping Zhong, Dinghao Zhou +2
Conformer has shown a great success in automatic speech recognition (ASR) on many public benchmarks. One of its crucial drawbacks is the quadratic time-space complexity with respec…