most citedStep-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

1 citations · 2 across the 6 of their papers we have counts for

collaborators

21 papers

cs.CL2026

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

Xiao Zhang, Qumeng Sun, Jiahao Li +4

Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy st…

cs.AI2026

AutoResearch: Insight In, Hallucination Out

Yiming Ren, Xiang Liu, Qumeng Sun +4

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically gr…

cs.SD2026

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

Qingjian Lin, Yuxin Li, Haoyang Zhang +14

Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale languag…

eess.AS2026

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Haoyang Zhang, Jun Chen, Donghang Wu +13

Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses.…

cs.SE2026

From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution

Junjie Wang, Yiming Ren, Haoyang Zhang

This beta technical report asks how reusable experience should be represented so that it can function as effective test-time control and as a substrate for iterative evolution. We…

eess.AS2026

Step-Audio-R1.5 Technical Report

Yuxin Zhang, Xiangyu Tony Zhang, Daijiao Liu +16

Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic…