30 citations · 42 across the 11 of their papers we have counts for
11 papers
EffVOC: Low-Delay Efficient Speech Waveform Reconstruction from Spectral Representations Without Phase
Renzheng Shi, Simon Welker, Timo Gerkmann +1
The Griffin-Lim algorithm has been a seminal contribution for phase reconstruction from amplitude spectrograms, however, requiring (infinitely) high algorithmic delay. Its low-dela…
Knowledge Distillation for Efficient Acoustic Echo Control
Ernst Seidel, Pejman Mowlaee, Tim Fingscheidt
In recent years, many efforts have been made to supersede classical acoustic echo control (AEC) algorithms with more powerful machine-learned approaches. While surpassing the perfo…
Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
Zhengyang Li, Thomas Graave, Björn Möller +3
In audiovisual automatic speech recognition (AV-ASR) systems, information fusion of visual features in a pre-trained ASR has been proven as a promising method to improve noise robu…
Engineering of Hallucination in Generative AI: It's not a Bug, it's a Feature
Tim Fingscheidt, Patrick Blumenberg, Björn Möller
Generative artificial intelligence (AI) is conquering our lives at lightning speed. Large language models such as ChatGPT answer our questions or write texts for us, large computer…
OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
Björn Möller, Zhengyang Li, Malte Stelzer +4
Recent successful video generation systems that predict and create realistic automotive driving scenes from short video inputs assign tokenization, future state prediction (world m…
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
Patrick Blumenberg, Thomas Graave, Tim Fingscheidt
Large language models (LLMs) demand extensive memory capacity during both fine-tuning and inference. To enable memory-efficient fine-tuning, existing methods apply block-wise quant…