activity
20242026
most citedKimi-Audio Technical Report

2 citations · 2 across the 6 of their papers we have counts for

collaborators

8 papers

cs.SD2026

UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization

Dongchao Yang, Yuanyuan Wang, Dading Chong +3

We study two foundational problems in audio language models: (1) how to design an audio tokenizer that can serve as an intermediate representation for both understanding and genera…

cs.CL2026

V-FAT: Benchmarking Visual Fidelity Against Text-bias

Ziteng Wang, Yujie He, Guanliang Li +3

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on standard visual reasoning benchmarks. However, there is growing concern…

cs.CV2025

MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning

Yueqian Wang, Songxiang Liu, Disong Wang +4

Recent advances in video multimodal large language models (Video MLLMs) have significantly enhanced video understanding and multi-modal interaction capabilities. While most existin…

cs.AI2025

Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning

Dongchao Yang, Songxiang Liu, Disong Wang +3

Recent advances in Omni models have enabled unified multimodal perception and generation. However, most existing systems still exhibit rigid reasoning behaviors, either overthinkin…

eess.AS20252 cited

Kimi-Audio Technical Report

KimiTeam, Ding Ding, Zeqian Ju +37

We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, inclu…

cs.SD2025

ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling

Dongchao Yang, Songxiang Liu, Haohan Guo +9

Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the ap…