activity
20242026
collaborators

9 papers

cs.CV2026

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

Qiming Li, Shujie Hu, Haohan Liu +3

Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks predominantly focus on single-t…

eess.AS2026

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu +1

Neural speech codecs efficiently compress speech and have become a foundation for speech generation, but they are typically learned as holistic representations that intertwine ling…

eess.AS2026

Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis

Zhikang Niu, Shujie Hu, Jeongsoo Choi +8

Mel-spectrograms have been widely used in zero-shot text-to-speech (TTS); their inherent redundancy leads to inefficiency in text-speech alignment. Compact VAE-based latent represe…

cs.SD2026

LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space

Detai Xin, Shujie Hu, Chengzuo Yang +4

We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that…

cs.CL2026

MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination

Zhuo Li, Yupeng Zhang, Pengyu Cheng +8

Hallucination remains a critical bottleneck for large language models (LLMs), undermining their reliability in real-world applications, especially in Retrieval-Augmented Generation…

cs.LG2025

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Haiyang Sun, Shujie Hu, Shujie Liu +8

Zero-shot streaming text-to-speech is an important research topic in human-computer interaction. Existing methods primarily use a lookahead mechanism, relying on future text to ach…