activity
20242026
collaborators

9 papers

cs.CV2026

VisualActBench: Can VLMs See and Act like a Human?

Daoan Zhang, Pai Liu, Xiaofei Zhou +6

Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely…

cs.SD2025

M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR

Ruixiang Mao, Xiangnan Ma, Qing Yang +7

The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mappin…

cs.CL2025

SUBQRAG: Sub-Question Driven Dynamic Graph RAG

Jiaoyang Li, Junhao Ruan, Shengwei Tang +5

Graph Retrieval-Augmented Generation (Graph RAG) effectively builds a knowledge graph (KG) to connect disparate facts across a large document corpus. However, this broad-view appro…

cs.CL2025

MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction

Jianjin Wang, Runsong Zhao, Xiaoqian Liu +6

Current direct speech-to-speech translation methods predominantly employ speech tokens as intermediate representations. However, a single speech token is not dense in semantics, so…

cs.CL2025

FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction

Yuan Ge, Saihan Chen, Jingqi Xiao +5

Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking…

cs.CL2025

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System

Yanfan Du, Jun Zhang, Bin Wang +6

Recent advances in speech large language models (SLMs) have improved speech recognition and translation in general domains, but accurately generating domain-specific terms or neolo…