most citedMini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models

1 citations · 1 across the 8 of their papers we have counts for

collaborators

9 papers

cs.CL2026

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Guibin Zhang, Leo Lu, Fangzhou Xie +15

Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can domi…

eess.AS2026

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Zhifei Xie, Jiaqi Lang, Ze An +7

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple…

cs.SD2026

MMAE: A Massive Multitask Audio Editing Benchmark

Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35

We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…

cs.SD2026

Audio Interaction Model

Zhifei Xie, Zihang Liu, Ze An +8

Audio is continuous and interactive, yet most Large Audio Language Models (LALMs) remain offline and streaming systems usually specialize in ASR or spoken dialogue. We formalize th…

cs.SD2026

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

Zhifei Xie, Kaiyu Pang, Haobin Zhang +4

Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustne…

cs.CL2026

Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation

Fangda Ye, Zhifei Xie, Yuxin Hu +5

Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centr…