53 citations · 181 across the 30 of their papers we have counts for
6 papers · 1 filter
Conditional Modeling Based Automatic Video Summarization
Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2
The aim of video summarization is to shorten videos automatically while retaining the key information necessary to convey the overall story. Video summarization methods mainly rely…
Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
Chao-Han Huck Yang, Yile Gu, Yi-Chieh Liu +3
We explore the ability of large language models (LLMs) to act as speech recognition post-processors that perform rescoring and error correction. Our first focus is on instruction p…
Can Whisper perform speech-based in-context learning?
Siyin Wang, Chao-Han Huck Yang, Ji Wu +1
This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SIC…
Causal Video Summarizer for Video Exploration
Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2
Recently, video summarization has been proposed as a method to help video exploration. However, traditional video summarization models only generate a fixed video summary which is…
Causalainer: Causal Explainer for Automatic Video Summarization
Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2
The goal of video summarization is to automatically shorten videos such that it conveys the overall story without losing relevant information. In many application scenarios, improp…
From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition
Chao-Han Huck Yang, Bo Li, Yu Zhang +4
In this work, we propose a new parameter-efficient learning framework based on neural model reprogramming for cross-lingual speech recognition, which can \textbf{re-purpose} well-t…