activity
20182023
most citedGenerative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting

53 citations · 181 across the 30 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

cs.CV2023★ 1 cited

Conditional Modeling Based Automatic Video Summarization

Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2

The aim of video summarization is to shorten videos automatically while retaining the key information necessary to convey the overall story. Video summarization methods mainly rely…

cs.CL2023★ 53 cited

Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting

Chao-Han Huck Yang, Yile Gu, Yi-Chieh Liu +3

We explore the ability of large language models (LLMs) to act as speech recognition post-processors that perform rescoring and error correction. Our first focus is on instruction p…

eess.AS2023★ 3 cited

Can Whisper perform speech-based in-context learning?

Siyin Wang, Chao-Han Huck Yang, Ji Wu +1

This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SIC…

cs.CV2023

Causal Video Summarizer for Video Exploration

Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2

Recently, video summarization has been proposed as a method to help video exploration. However, traditional video summarization models only generate a fixed video summary which is…

cs.CV2023

Causalainer: Causal Explainer for Automatic Video Summarization

Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2

The goal of video summarization is to automatically shorten videos such that it conveys the overall story without losing relevant information. In many application scenarios, improp…

cs.SD2023★ 26 cited

From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition

Chao-Han Huck Yang, Bo Li, Yu Zhang +4

In this work, we propose a new parameter-efficient learning framework based on neural model reprogramming for cross-lingual speech recognition, which can \textbf{re-purpose} well-t…