activity
20242026
most citedZ-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

1 citations · 1 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CV20261 cited

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Image Team, Huanqia Cai, Sihan Cao +21

The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. Leading open-source alternatives,…

cs.CL2026

JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR

Shilin Zhou, Zhenghua Li

Contextual Automatic Speech Recognition (ASR) faces challenges with large-scale keyword dictionaries, as excessive irrelevant candidates introduce noise that degrades accuracy. To…

cs.LG2025

Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models

Zhichao Sheng, Shilin Zhou, Chen Gong +1

Large Audio Language Models (LALMs), powered by the chain-of-thought (CoT) paradigm, have shown remarkable reasoning capabilities. Intuitively, different problems often require var…

cs.MM2025

Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision

Che Liu, Yingji Zhang, Dong Zhang +13

This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limit…

eess.AS2025

UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets

Zhichao Sheng, Shilin Zhou, Chen Gong +1

Spoken Language Understanding (SLU) plays a crucial role in speech-centric multimedia applications, enabling machines to comprehend spoken language in scenarios such as meetings, i…

cs.CL2025

Improving Contextual ASR via Multi-grained Fusion with Large Language Models

Shilin Zhou, Zhenghua Li

While end-to-end Automatic Speech Recognition (ASR) models have shown impressive performance in transcribing general speech, they often struggle to accurately recognize contextuall…