1 citations · 1 across the 2 of their papers we have counts for
7 papers
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Image Team, Huanqia Cai, Sihan Cao +21
The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. Leading open-source alternatives,…
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
Shilin Zhou, Zhenghua Li
Contextual Automatic Speech Recognition (ASR) faces challenges with large-scale keyword dictionaries, as excessive irrelevant candidates introduce noise that degrades accuracy. To…
Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models
Zhichao Sheng, Shilin Zhou, Chen Gong +1
Large Audio Language Models (LALMs), powered by the chain-of-thought (CoT) paradigm, have shown remarkable reasoning capabilities. Intuitively, different problems often require var…
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
Che Liu, Yingji Zhang, Dong Zhang +13
This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limit…
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
Zhichao Sheng, Shilin Zhou, Chen Gong +1
Spoken Language Understanding (SLU) plays a crucial role in speech-centric multimedia applications, enabling machines to comprehend spoken language in scenarios such as meetings, i…
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
Shilin Zhou, Zhenghua Li
While end-to-end Automatic Speech Recognition (ASR) models have shown impressive performance in transcribing general speech, they often struggle to accurately recognize contextuall…