works on

From the 1 of 15 linked papers with an AI index.

activity
20242026
collaborators

15 papers

cs.LG2026

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

Kaiyang Ye, Yuan Ge, Junxiang Zhang +8

While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored.…

cs.CL2026

ToFu: A White-Box, Token-Efficient Agent Harness for Researchers

Junhao Ruan, Yuan Ge, Bei Li +7

ToFu is an open‑source, white‑box agentic harness that lets researchers automate codebase reading, file editing, command execution, and tool integration with high token efficiency…

cs.SD2026

DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast

Zhengkun Ge, Xiaoqian Liu, Haoran Zhang +5

Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training-free methods typically rely o…

cs.CL2026

On the Emotion Understanding of Synthesized Speech

Yuan Ge, Haishu Zhao, Aokai Hao +10

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesize…

cs.CL2026

StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control

Haishu Zhao, Aokai Hao, Yuan Ge +3

Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For mor…

cs.SD2026

When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning

Ruixiang Mao, Xiangnan Ma, Dan Chen +12

Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive p…