◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Florian Metze

14 papers hereh-index 6253 citations21 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author6
  • last author6

Across the 12 of 14 papers where every author was matched, so the position is known.

fields
  • eess.AS6
  • cs.CL5
  • cs.AI1
  • cs.CV1
  • cs.SD1
same name
  • Florian Metze — 72 papers, h 45
  • Florian Metze — 1 paper
  • Florian Metze — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20172026
most citedEmbodied AI Agents: Modeling the World

9 citations · 9 across the 13 of their papers we have counts for

collaborators
Showing 2026Show all

4 papers · 1 filter

eess.AS2026

Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning

Zhicheng Ouyang, Seong-Gyun Leem, Bach Viet Do +4

Conversational AI has made significant progress, yet generating expressive and controllable text-to-speech (TTS) remains challenging. Specifically, controlling fine-grained voice s…

cs.CL2026

Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning

Jingxiang Chen, Minseok Kim, Seong-Gyun Leem +13

Speech large language models (LLMs) observe paralinguistic cues such as prosody, emotion, and non-verbal sounds--crucial for intent understanding. However, leveraging these cues fa…

cs.CL2026

Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities

Ju Lin, Jing Pan, Ruizhi Li +5

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech understanding capabilities. However, most speech LLMs are…

eess.AS2026

T-Mimi: A Transformer-based Mimi Decoder for Real-Time On-Phone TTS

Haibin Wu, Bach Viet Do, Naveen Suda +10

Neural audio codecs provide promising acoustic features for speech synthesis, with representative streaming codecs like Mimi providing high-quality acoustic features for real-time…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.