9 papers
VisualActBench: Can VLMs See and Act like a Human?
Daoan Zhang, Pai Liu, Xiaofei Zhou +6
Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely…
M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
Ruixiang Mao, Xiangnan Ma, Qing Yang +7
The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mappin…
SUBQRAG: Sub-Question Driven Dynamic Graph RAG
Jiaoyang Li, Junhao Ruan, Shengwei Tang +5
Graph Retrieval-Augmented Generation (Graph RAG) effectively builds a knowledge graph (KG) to connect disparate facts across a large document corpus. However, this broad-view appro…
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
Jianjin Wang, Runsong Zhao, Xiaoqian Liu +6
Current direct speech-to-speech translation methods predominantly employ speech tokens as intermediate representations. However, a single speech token is not dense in semantics, so…
FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction
Yuan Ge, Saihan Chen, Jingqi Xiao +5
Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking…
Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System
Yanfan Du, Jun Zhang, Bin Wang +6
Recent advances in speech large language models (SLMs) have improved speech recognition and translation in general domains, but accurately generating domain-specific terms or neolo…