◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Zhuo Chen

4 papers hereh-index 468 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • eess.AS2
  • cs.CL1
  • cs.SD1
same name
  • Zhuo Chen — 49 papers, h 44
  • Zhuo Chen — 27 papers, h 23
  • Zhuo Chen — 15 papers, h 5
  • Zhuo Chen — 13 papers, h 6
  • Zhuo Chen — 11 papers, h 5
  • Zhuo Chen — 10 papers, h 6

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232025
most citedICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge

1 citations · 1 across the 3 of their papers we have counts for

collaborators

4 papers

eess.AS2025

SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning

Prabhat Pandey, Rupak Vignesh Swaminathan, K V Vijay Girish +4

We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50…

cs.SD2024★ 1 cited

ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge

He Wang, Pengcheng Guo, Yue Li +13

To promote speech processing and recognition research in driving scenarios, we build on the success of the Intelligent Cockpit Speech Recognition Challenge (ICSRC) held at ISCSLP 2…

cs.CL2023

COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Jing Pan, Jian Wu, Yashesh Gaur +4

We present a cost-effective method to integrate speech into a large language model (LLM), resulting in a Contextual Speech Model with Instruction-following/in-context-learning Capa…

eess.AS2023

t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability

Jian Wu, Naoyuki Kanda, Takuya Yoshioka +3

Token-level serialized output training (t-SOT) was recently proposed to address the challenge of streaming multi-talker automatic speech recognition (ASR). T-SOT effectively handle…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.